From the very beginning, video games have been closely intertwined with the development of computer science. Over the past few decades, breakthroughs in real-time rendering and hardware have pushed games toward richer visuals, larger worlds, and more capable engines.

Yet the characters at the heart of the player experience have not advanced at the same pace. It remains one of the game industry's most persistent problems: making characters feel genuinely lifelike.

In the real world, the meaning of a journey is never determined by the destination, but by the people who travel with you. The same is true for virtual worlds. True immersion does not come from visual fidelity, but from the bond between you and these characters. They remember what you said, understand what you did, and in response, change their own fate and shape your story.

In traditional games, NPCs are driven by extensive hand-authored scripts and branching behavior trees, so the experience is largely pre-authored. This paradigm can deliver stable, controllable gameplay in the short term, but it has an inherent limit. As players spend more time and repeat the same kinds of interactions, novelty fades quickly, characters become predictable, the world feels less responsive, and players eventually get bored and leave.

For example, in The Elder Scrolls V: Skyrim, guards are designed to trigger different lines based on the player's game state, but players soon find themselves hearing the same familiar line again and again: "I used to be an adventurer like you. Then I took an arrow in the knee…" This frequent repetition eventually became a well-known gaming meme.

At Entropy, we built AI models and infrastructure that power the next generation of games: systems that enable characters to have persistent memory, stable personalities, intent-aware conversation, meaningful actions, and natural, lifelike voices.

These characters treat your pauses, pacing, choices—and even silence—as signals. They read and remember those cues, shaping what they say and do next. The story evolves through ongoing interaction rather than fixed branches. The result is a game that feels genuinely personal.

To get there, we need to bring frontier AI into games and integrate it tightly with the game world. That introduces several core challenges.

  • The biggest challenge is cost. High inference costs make broad adoption difficult. While per-token pricing has fallen rapidly over the past few years, it is still far from a level that commercial games can sustainably support. Today, even the most affordable models can cost anywhere from a few dollars to tens of dollars for several hours of gameplay. This cost may exceed the purchase price of many games, making it economically impractical to bring advanced AI into games.
  • Latency is the second bottleneck. In a cloud inference architecture, every player request must be processed server-side and returned to the client, so responsiveness depends on network latency and connection stability. AI-driven logic forces players to wait several seconds, even though rendering and physics simulation run locally in real time. This puts the AI a few beats behind the rest of the game. NPC dialogue and actions arrive a few seconds late, making characters feel sluggish and out of sync.
  • Character consistency and out-of-character (OOC) drift. Today's AI models are designed for general-purpose conversation rather than for games. As a result, they lack a reliable sense of diegetic boundaries.

    When asked to role-play as in-game characters, they are prone to drifting beyond the game's narrative constraints and the character's established background. A "medieval physician," for example, may suggest buying amoxicillin at a pharmacy or using Google Maps to locate the nearest hospital. Such breakdowns quickly undermine player trust and can shatter immersion.

  • Safety and hallucination risks. Large language model (LLM) hallucinations are especially damaging in games. The model may fabricate lore or narrative events that fall outside the established worldbuilding. What's worse, it may generate extreme content—hate, discrimination, explicit material, or graphic violence—leaving players deeply uncomfortable.
  • Speech models lack dialogue awareness. Even the most advanced systems still operate within a text-to-speech (TTS) paradigm: they convert written text into speech waveforms and primarily optimize for acoustic fidelity. In doing so, they largely overlook the implicit structure and cues that shape human dialogue, including turn-taking, interruptions, and delivery signals like timing, pacing, and tone.

    Game dialogue places even greater demands on speech generation. An NPC's spoken response must not only reflect the player's emotion and intent but also align with the current in-game situation. Yet traditional TTS pipelines remain largely disconnected from dialogue context and game state, leaving NPC speech far from natural or lifelike.

  • Traditional behavior trees lack native support for LLM-driven deliberation. They pass around control-flow states—success, failure, running, aborted—but those signals never enter the language model's cognitive loop. As a result, the model may produce plausible dialogue without a reliable mechanism to drive NPC action selection or execution. An NPC might claim it has launched a magic orb, but the game requires concrete outcomes: a casting animation, in-world effects, and changes to the game state.

    A similar disconnect exists at the narrative layer. While dialogue may be open-ended, the underlying plot often remains hard-locked to pre-authored story nodes. That leaves characters unable to use conversation and interaction history to advance the storyline in a dynamic and responsive way.

These challenges point to a simple conclusion: building an AI-native game is not primarily about using a smarter model. It is about engineering an end-to-end system that is cost-efficient, low-latency, controllable, and safe by design.

That system should be centered on the player experience and tightly integrated with core gameplay. AI games only become genuinely fun and scalable when AI moves beyond surface-level dialogue and becomes a reliable part of the narrative-and-action loop.

We begin by introducing Entropy's system architecture and show, step by step, how it solves the six most critical challenges of deploying AI in games. Next, we explain why a local-first approach is the most practical way to make AI characters fast, affordable, and sustainable in real gameplay. Finally, we cover the technical core: model and architecture decisions, performance trade-offs, and the system-level techniques required to ship AI in production games.

How NPCs think and stories evolve

Inspired by cognitive neuroscience, we built a coordinated, multi-module cognitive system for NPCs.

The system centers on a language model (LM) as its cognitive core. It runs in a closed loop with modules for perception, memory, context engineering, behavior control, and speech. At a higher level, an AI Narrative Director manages narrative pacing and player experience, turning hours of interaction into a coherent, controllable story.

Click any card in the diagram to highlight its path; click any brain module for a brief description.

In this framework, NPCs rely on a perception layer to take in both the in-world state and the player's language and actions. The memory system is split into short-term memory (STM) for immediate decisions and long-term memory (LTM) for longer-horizon decisions and stable character consistency.

We trained and deployed a Game-Native Language Model (GNLM) as the NPC's cognitive core. It maintains a stateful cognition loop that tracks the evolving game state and interaction history, enabling NPCs to adjust tone, response strategy, and actions in real time.

The behavior and cognition layers do not form a one-way decide-then-execute pipeline. Instead, they operate as a bidirectional feedback loop. The cognition layer issues action and speech directives, while the behavior layer continuously streams back execution state, outcomes, and in-world signals. Those signals immediately shape the next round of deliberation and generation, enabling iterative correction in real time.

The diagram above represents the gameplay loop as a graph of narrative nodes and paths. Each player–NPC interaction—dialogue, actions, shared events, and the resulting memory updates—directly shapes the NPC's subsequent decisions and behavior policy. The outcomes of the NPC's speech and actions then flow back into the cognitive modules and help the AI Narrative Director steer the narrative.

In our architecture, the AI Narrative Director functions as a player-experience management layer. It continuously monitors the player's game state and key player-experience signals, including objective progress, resource pressure, event density, and immersion cues.

Based on those signals, it adjusts event triggers, information reveals, conflict intensity, and narrative pacing without breaking world consistency or character motivation. The aim is to sustain long-term novelty and challenge within clear, designer-defined boundaries, so that different players can develop distinct stories under the same world rules.

How do we solve the toughest challenges and bring advanced AI into game worlds?

We now outline six core problems in game AI and explain how our system addresses them. This work is ongoing and, in our view, marks the start of a new paradigm for games. Over time, these technical directions and implementations will fundamentally reshape how games are built and played.

  • We run inference locally rather than in the cloud, cutting per-session inference costs from tens of dollars to effectively zero. Players pay no API fees tied to playtime.

    To make on-device deployment the default, we built a dedicated inference runtime that hosts the language model, speech model, and a set of lightweight specialized models. The runtime is frame-aware and runs as part of the game loop. It coordinates execution and schedules inference within the frame-time budget to preserve smooth rendering and simulation.

    We further improve efficiency with quantization, mixed precision, and knowledge distillation to ensure reliable performance across consumer hardware. In addition, we integrate each model and its inference stack directly into the game runtime so the AI stays invisible to the player. In practice, this approach runs on NVIDIA RTX 30-, 40-, and 50-series GPUs and Apple Silicon Macs (M1–M4).

  • The entire stack is streaming-first by default, so NPC interactions stay responsive in real time. Players hear the first words of a response without waiting for the language or speech model to complete a full sentence. Game-specific models are kept lightweight enough to run in parallel. Together, these choices reduce end-to-end conversational latency to approximately 500 ms.

    We are also exploring a full-duplex interaction mode. Games come with different audio and interaction constraints than voice assistants, especially in noisy and dynamic environments. For that reason, we plan to introduce full-duplex capabilities gradually and with care.

  • Character consistency requires targeted intervention rather than relying on a single training pass. We run multiple rounds of post-training tailored to game-specific requirements. These rounds enforce diegetic boundaries, improve memory reliability, handle pronouns and perspective correctly, and adhere strictly to developer-defined lore and narrative worlds. An internal evaluation pipeline measures these capabilities and tracks the best-performing checkpoints.

    Context management is a separate but closely related challenge. We address it by training a game-specific retrieval model and using a dedicated memory module to manage both long-term and short-term NPC memory. Together, these components form our context-engineering stack. It keeps runtime context efficient, accurate, and genuinely useful during gameplay.

    In practice, a stronger base model combined with effective context engineering can reduce lore violations and out-of-character behavior. However, it remains an open research problem in machine learning to make a model faithfully and consistently embody a developer-authored character. To push beyond current limits, we're starting interpretability work to study activation patterns across story branches and interaction contexts, so we can better understand what keeps a character in character.

  • Safety is our top priority. We treat safety as a design constraint, not an afterthought. We explicitly incorporate major risk categories, including hate, discrimination, explicit content, and violence, into our post-training pipeline. This reduces the likelihood that the model learns or reproduces harmful behaviors during play.

    Looking ahead, we plan to integrate a dedicated safeguard model, similar in spirit to gpt-oss-safeguard. It screens high-risk cases before they reach the player, ensuring a consistently safe in-game experience.

  • We built our speech model from scratch so NPC voice interactions feel natural and lifelike. The model is designed to capture conversational nuance and in-game context. We also engineered our training pipeline and optimization objectives for both pre- and post-training to enable high-quality, context-aware speech at ultra-low latency.

    This approach goes beyond traditional text-to-speech systems that synthesize audio from text alone. Given the game setting and the ongoing conversation, the model can adjust emotion, tone, and delivery to match both the moment and the character, creating a more immersive in-game dialogue experience.

    For us, this is only the beginning. We are continuing to advance this work so the model can understand not only what players say, but also how they say it, including timing, speaking style, emotion, and engagement cues, in real time and at low latency. This helps the system better understand the player and deliver the most appropriate response for the situation.

  • We extend the traditional Behavior Tree (BT) into a Cognitive Tree (CT) that reasons over natural-language semantics. Execution state and player behavior signals—blocking, turning away, approaching, silence, and more—feed back into the AI's cognitive system as semantic inputs. NPCs then adapt not only to what was said, but also to what those in-world signals imply. Is the player's silence hesitation or refusal? Does following closely indicate trust, or a need for guidance? This bidirectional loop keeps language and action aligned.

    For example, when an NPC says, "I'm going to cast a fireball," the game executes the cast and the player sees the fireball appear in the world. If the player interrupts the cast, the NPC can infer "the player doesn't want me to do this" and adjust its subsequent dialogue and behavior.

    At the narrative level, the AI Narrative Director continuously monitors and aggregates player–NPC interaction and dialogue history. When it detects a drop in player engagement, it can inject new events and unlock additional story branches. This shifts the game's narrative logic so each player's story can unfold along a distinct trajectory.

LLM Inference Cost

$/1M tokens
Claude Sonnet 4.5$9.00
GPT-5.2$7.88
Gemini 3 Pro$7.00
Gemini 3 Flash$1.75
DeepSeek V3$0.35
Entropy (Local)FREE

Cost per 3h Interactive Session

TTS Cost

$/1M characters
ElevenLabs v3$206.00
Google Studio$160.00
MiniMax Speech 2.6 Turbo$60.00
Azure Neural$15.00
Inworld TTS 1 Max$10.00
Entropy (Local)FREE

Cost Efficiency. (Left) LLM inference unit pricing (USD per 1M tokens; blended input/output). (Center) Cumulative end-to-end API cost over a 3‑hour gameplay session (≈460 dialogue turns) in our sandbox testbed, computed from measured token/character usage and the corresponding provider list prices; voice via ElevenLabs v3. (Right) TTS unit pricing (USD per 1M characters). Prices are sourced from official provider documentation (Jan 2025). Entropy runs fully on-device and therefore incurs zero marginal API cost.

LLM inference cost in USD per one million tokens
ProviderCost
Claude Sonnet 4.59
GPT-5.27.88
Gemini 3 Pro7
Gemini 3 Flash1.75
DeepSeek V30.35
Entropy (Local)Free
Three-hour interactive session cost in USD
ProviderThree-hour session cost
Sonnet-4.533.32
Gemini-3 Pro30.2
GPT-5.229.63
Gemini-3 Flash25.17
DeepSeek-v324.24
Entropy0
TTS cost in USD per one million characters
ProviderCost
ElevenLabs v3206
Google Studio160
MiniMax Speech 2.6 Turbo60
Azure Neural15
Inworld TTS 1 Max10
Entropy (Local)Free

Why local-first is the best solution for game AI

The fundamental flaw of an API-based cloud inference model is that it breaks the core economics of games: the longer players play, the more developers should benefit.

In many traditional games, marginal costs are close to zero. Whether a player puts in 100 hours or 1,000, server costs typically do not rise in proportion to playtime. This lets developers focus on what matters most: making the game more fun. As players spend more time, the community grows and more content gets created. That content attracts new players, creating a positive flywheel.

However, cloud-based AI inverts this logic. Each line of dialogue and each NPC reaction triggers a paid API call, so inference spending scales with interaction volume and session length. The more a player engages, the more the developer pays. In practice, this leaves developers with two options: meter AI usage or impose hard caps on AI-driven behavior. Both approaches degrade the player experience.

Local-first inference restores the natural economics of games. Marginal costs are effectively zero, so players can play as long as they want. Developers can use the same business models that work in traditional games—one-time purchase, subscription, or free-to-play with in-game purchases. They are no longer constrained by a cloud cost structure that rises with engagement.

With a local-first approach, we reduce both friction and the recurring inference cost of game AI to near zero. We keep the focus where it belongs: building a great game. Players buy the game, not tokens.

Local AI Game FlywheelCloud AI Game SpiralBetter gameMore playtimeMore revenueRising inference costMore playersCaps & paywallsCommunity & UGC growthWorse experienceBetter experienceLower retentionUnrestricted interactionLess revenueZero marginal costWorse gameMore playtimeChurn

Beyond cost, local-first offers practical advantages that are easy to miss.

  • For players, on-device AI means the gameplay experience no longer requires permission from a remote service. There is no dependence on network stability, no API rate limiting, and no conversations cut short by timeouts. Whether a player is on a long-haul flight or in a family cabin, the full AI system can keep running.

  • Local games can offer something cloud games struggle to guarantee: permanence and ownership. Every year, long-running online games shut down their servers, and thousands of hours of play, communities, and shared memories can disappear overnight. A local-first architecture changes that dynamic. The game's lifespan is determined by players and the community, not by a company's operational decisions.

  • It also makes privacy the default. Dialogue and voice data stay on the player's device, without requiring them to accept mandatory data-collection terms.

    All of this is seamless for players. There is no separate model download and no API key configuration. Developers ship the model weights and inference runtime with the game at release, making the system plug-and-play from day one.

  • Local-first can unlock a larger creator ecosystem. A game's fun comes not only from the content it ships with, but also from the creator community that grows around it. With the right UGC tools, developers and players can create surprising new things. Traditional games already show this compounding effect. Minecraft has evolved into a broad universe of worlds and playstyles, and Roblox hosts millions of creator-built experiences across a long tail of niche genres.

    Local-first AI games have a stronger foundation for reaching comparable scale. The key advantage is cost: near-zero marginal inference cost makes creation and experimentation cheap. At the same time, modifiable weights make models easier to remix, extend, and iterate. Over time, that combination can expand the range of gameplay and creative expression.

A longer-term trend is also becoming clear. Parameter efficiency continues to improve. On many tasks, today's ~1B-parameter models can match or even surpass earlier 100B-scale models. This makes models that are smart enough yet small enough to run on-device increasingly practical.

In parallel, more hardware companies are prioritizing inference, and next-generation GPUs and NPUs are being optimized at the system level for inference workloads. With each generation, consumer GPUs keep raising the ceiling on what can run locally. Together, these shifts will continue to drive down the unit cost of on-device inference, moving local-first AI from a technical preference to the most economical and natural default for games.

To borrow Andrej Karpathy's analogy, the AI industry today resembles the mainframe era. A small number of companies operate massive GPU clusters and distribute AI capability to everyone through APIs. In the early days of computing, only large institutions, including governments, universities, and major companies, could afford computers, and everyone else had to share access to compute time. But this is not the end state.

As smaller models become more capable and consumer hardware continues to advance, demand will shift from a Model-as-a-Service paradigm to an Application-First paradigm built around specific products. This will move us into AI's personal computing era, where intelligence no longer lives only in the cloud but runs on your own device. It understands you, remembers you, and belongs to you.