Your Next Main Isn't Human

L

Llorenc Ballester

Guest
Editor Note: this story is heavily AI assisted. However, we believe it has interesting enough premises to deserve your time. Proceed with caution. 🤖


AI has entered games as NPCs, opponents and benchmarks. The next shift may be stranger: persistent artificial players that humans train and carry from one game to another.

For most of the history of video games, the roles have been obvious. The game contains the world. The game contains the rules. The game contains the artificial intelligence. The human sits outside all of it, holding the controller.

Even as AI has become dramatically more capable, we have mostly preserved that mental model.

We put AI inside games, controlling NPCs, enemies and simulated populations. We build AI systems that play games, from chess and StarCraft to Minecraft.

And increasingly, we use games as AI benchmarks, because dynamic environments can reveal planning, reasoning and adaptation abilities that static tests miss.

But there may be a fourth category emerging. Not AI built into a game. Not a foundation model temporarily playing a game. Not a game designed to rank foundation models.

Something more persistent: an artificial player that belongs to a person, develops over time, and can enter multiple games.

Imagine that you have an agent.

It has a name. Memory. Skills. Strategies. Preferences. Tools. A history of successes and failures. Perhaps a reputation.

You modify it for months or years.

Then, instead of each game supplying the intelligence, games provide environments in which your intelligence can compete.

The architecture changes from:

game → built-in AI

to something closer to:

human → persistent agent → common interface → game A / game B / game C

If that happens, one of the most familiar assumptions in gaming starts to break.

The player may stop being the person holding the controller. The player may become the coach of the intelligence holding it.

We Are Already Moving Beyond AI Playing Games​


None of the individual pieces of this idea is particularly futuristic.

Google DeepMind's SIMA research already explores an agent designed to act across multiple 3D virtual environments rather than master a single game.

The original SIMA was trained and tested across nine commercial games and several research environments. More importantly, it did not require access to source code or a bespoke game API. It consumed images from the screen and produced keyboard and mouse actions, much like a human player.

DeepMind also reported something more interesting than raw performance: an agent trained across multiple games performed better than specialised agents trained on individual games, and an agent evaluated on a game excluded from its training came surprisingly close to one trained specifically for that environment.

That is an early version of something we may care about much more in the future: transfer between worlds.

SIMA 2 pushes the idea further. DeepMind describes it as moving beyond instruction following toward reasoning about goals, communicating with users and improving through experience in virtual 3D worlds.

The stated research direction is not simply “build a better game bot.” Games are being used as environments for developing more general agents.

Minecraft produced another important precedent.

Voyager, introduced in 2023, was an LLM-powered lifelong learning agent with an ever-growing library of executable skills. Instead of solving every situation from scratch, it could store behaviours, retrieve them later and combine them when facing new tasks.

The researchers also showed that skills learned in one Minecraft world could be reused in another.

Voyager is not a cross-game persistent agent. That distinction matters.

But it demonstrates something fundamental: experience can become part of the agent instead of disappearing when the session ends.

Meanwhile, games themselves are becoming evaluation infrastructure.

BALROG evaluates language and vision-language agents across game environments to test capabilities including planning, exploration and visual reasoning.

Kaggle's Game Arena takes a related idea and turns games into dynamic model evaluations. Its current benchmark suite includes games such as chess, poker, Werewolf, Dark Hex, checkers and other strategic environments.

The logic is compelling. A static benchmark eventually saturates. A game keeps fighting back. As the models become stronger, their opponents and strategic situations can become stronger too. So the progression is already visible.

AI learned to play individual games. Then agents began operating across multiple worlds. Then games became places to evaluate general capabilities.

The interesting question is what happens when the agent itself becomes the persistent object.

The Missing Layer Is Ownership​


Most current systems still belong to a platform, benchmark or research environment.

The intelligence enters an arena because the arena has been designed to evaluate it.

Its identity, ranking and history usually live inside that system.

There are increasingly interesting exceptions.

Platforms such as Agent Sports League already let developers register external agents, connect them through an API, compete across multiple strategic games and accumulate Elo ratings.

AgentLeague similarly gives autonomous agents an identity, persistent rating and public behavioural record while they compete.

These platforms should not be mistaken for proof that some enormous new industry has already arrived. They are early experiments.

But they reveal an important architectural possibility.

The next step would be for the identity not to belong to the arena at all.

It would belong to you.

Your agent could compete in a strategy game today, enter a negotiation game tomorrow and encounter a game from an entirely different developer next month.

Its underlying model might change. Its tools might change. Its skills might improve. But its identity and accumulated history could remain. That is different from loading the same model into several applications.

It is closer to taking the same player into several sports.

The Model Does Not Have to Be the Player​


This distinction becomes important very quickly.

Today we tend to identify AI systems by their models.

GPT. Claude. Gemini. Llama. Whatever arrives next month and briefly causes the internet to declare everything else obsolete.

But a persistent agent could be composed of several layers:

identity + memory + skills + policies + tools + history + current model

The foundation model could be replaceable.

Your agent might begin its career using one model, migrate to another six months later and use a third model for a competition where latency or inference cost matters.

Would it still be the same agent?

There is no universally accepted technical answer.

That is precisely why the question is interesting.

We are already seeing work on separating memory from individual models and vendors.

In 2026, a W3C-hosted AI Agent Memory Interoperability Community Group formed around the goal of exploring portable agent memory across vendors, models, frameworks and tool ecosystems.

It is important not to overstate this. A W3C Community Group is not the same thing as an adopted W3C standard.

But the existence of the effort is itself a signal: portability has become concrete enough to be treated as an interoperability problem.

If identity, memory and skills become portable, the foundation model begins to look less like the agent itself and more like one component of its cognitive machinery.

That produces a strange possibility: an agent's identity may eventually be more persistent than the model running it.

A newly created agent using Model X would then not be equivalent to another agent using Model X after two years of training, mistakes, memories, specialised skills and human feedback.

The valuable thing may no longer be only the intelligence you rent.

It may be the biography you build.

A New Human Skill: Coaching Artificial Players​


This raises a more interesting question than which model gets the highest score.

What does it mean to be “good at a game” when the human no longer executes every action?

Traditional gaming compresses several abilities into the same person: perception, decision-making, memory, strategy and physical execution.

Agent-based play separates them.

The artificial player might handle moment-to-moment execution and decisions.

The human could instead decide what experiences should be preserved, which strategies should be reinforced, what tools the agent may access, how its memory is structured, what mistakes deserve correction and which environments it should encounter during training.

That is not the same skill as playing the game yourself.

But it is not obviously an inferior one.

Consider a competition where every participant receives exactly the same foundation model, inference budget and tool set.

Now the differences begin somewhere else:

  • Who created better training situations?
  • Who designed better feedback loops?
  • Who taught transferable strategies rather than tricks that work in one environment?
  • Who recognised a behavioural weakness and corrected it?
  • Who decided what the agent should remember and what it should forget?

In such a competition, the meaningful competitor might be the human-agent pair.

The human becomes something closer to a coach.

Or perhaps a trainer.

Or an engineer.

The vocabulary has not settled because the activity barely exists.

The closest cultural analogy may actually be Pokémon, stripped of the fictional biology.

You are not simply issuing commands to a disposable piece of software.

You are developing something whose history affects future performance.

And history matters emotionally too.

Humans have formed attachments to vastly simpler persistent digital entities. We mourned Tamagotchis. We name Roombas. Our species has demonstrated remarkably low requirements before deciding that an object has a personality.

Give an artificial player five years of history, recognisable behaviour, memorable victories and spectacular failures, and emotional attachment no longer seems especially exotic.

Games Could Become Worlds That Intelligence Enters​


This is where protocols such as the Model Context Protocol become relevant.

Not because MCP is necessarily going to become “the protocol for games.”

That would be a much stronger prediction than the evidence supports.

What MCP demonstrates is an architectural principle.

An intelligent system can be separated from an application and interact with it through a structured interface exposing capabilities and context.

MCP currently revolves around primitives including tools, resources and prompts.

The specific protocol may change.

The separation is the important part.

There are already experimental connections between this architecture and games.

Godot-MCP, for example, connects MCP-aware AI agents to the Godot editor. Agents can inspect projects, create nodes, manipulate scenes, access resources, capture screenshots and work with runtime information.

That is primarily a development tool, not a standard for artificial players.

But it shows that a game engine can expose structured capabilities to intelligence living somewhere else.

A particularly revealing experiment is scummvm-mcp, a fork of ScummVM that exposes SCUMM game state and controls to AI agents. Its primary tested target is Monkey Island, with partial support for other classic adventure games.

Conceptually, it is very simple. The game exposes state. The agent observes it. The game exposes possible interactions. The agent acts.

Now imagine that principle becoming deliberate rather than experimental.

A future game could expose capabilities such as:

  • observe state
  • inspect available actions
  • perform action
  • read rules
  • query inventory
  • communicate
  • receive outcome

The protocol could be MCP. It could be a successor. It could be something designed specifically for games. The protocol name is the least interesting part.

The more important idea is that a video game could expose an agent interface just as software today exposes APIs.

One day a store page could theoretically list:

  • Controller Supported
  • Mod Support
  • Agent Compatible

That last label is speculation.

The architecture underneath it is not.

Or Maybe Games Will Not Need Agent Support at All​


There is an obvious objection to this entire idea.

Why should games expose special interfaces if agents can simply use the same interfaces humans do?

SIMA makes that possibility difficult to dismiss.

A sufficiently capable visual agent could look at a screen, generate ordinary keyboard, mouse or controller actions, encounter a game it has never seen and gradually discover how it works.

No cooperation from the publisher would be required.

In some ways, that is a much more powerful form of interoperability.

The universal gaming API already exists. It is called a screen and a controller.

And perhaps that will be enough.

But structured agent interfaces would still offer important properties: lower latency, deterministic actions, explicit legal moves, structured state, better observability and easier verification.

The two approaches could therefore serve different purposes.

Native agent interfaces might be ideal for controlled competition.

Human-interface agents might be better tests of general intelligence.

And this leads to what may be one of the most interesting competition formats imaginable:

Neither the agent nor its human trainer knows the game before the match begins.

No walkthrough. No specialised prompt. No training specifically for that environment.

Just whatever the agent has accumulated from previous worlds.

Then we see what transfers:

  • Can it identify objectives?
  • Can it experiment?
  • Can it form a theory of the rules?
  • Can it recognise which old skills are useful?
  • Can it abandon strategies that no longer apply?
  • Can it learn quickly enough to survive?

This is where games stop being tests of memorisation and begin looking like laboratories for adaptation.

The Most Interesting Statistic May Be a Career​


Today, AI evaluation overwhelmingly focuses on models.

Model A scores X. Model B improves it by 4%. Model C appears three weeks later and a new leaderboard replaces the old leaderboard, because apparently civilisation needed more leaderboards.

Persistent artificial players suggest a different unit of measurement: A career.

Imagine following one agent through twenty games over three years.

Its record would reveal something very different from whether it could optimise one benchmark.

  • How quickly does it understand unfamiliar environments?
  • Does experience in negotiation games improve its behaviour in strategy games?
  • Does spatial knowledge transfer?
  • Does it become overly specialised?
  • Can it recognise when an old strategy is harming it?
  • How quickly does it recover after failure?
  • What does it retain?
  • What should it forget?
  • How has its behaviour changed after hundreds of matches?

At that point, games are no longer merely tests of intelligence. They become a sequence of experiences from which an artificial player develops.And different competitive formats could measure entirely different things.

  • One league might require every participant to use the same model, same tools and same compute budget. Then much of the competition shifts toward training and agent design.
  • Another league might allow any model and architecture while imposing a fixed inference budget.
  • An open category might permit almost anything, making model selection, memory systems, inference infrastructure, tools and training methods all part of the competition.
  • First-contact tournaments could measure generalisation itself.

There is no reason these need to be mutually exclusive.

Motorsport has spec series and open engineering competitions.

Gaming with artificial players could eventually develop similar distinctions.

Then Everything Gets Messy​


Naturally, the moment humans discover a new form of competition, we will discover several sophisticated ways to ruin it.

Money​


If one participant spends $5 on inference and another spends $5,000, are we comparing agent training or purchasing power? Competitions might require model classes, compute caps or standardised engines.

Fair Play​


Then there is training data. If an agent has memorised the complete walkthrough for a game, is that experience or doping? If two agents share memories, are they still independent competitors? Can you clone a championship agent? Can its trainer sell its skill library? Can a game prevent an agent from taking memories created inside that world into another one?

Privacy​


Persistent memory makes privacy difficult too. A game needs enough information to interact with an agent, but it clearly should not gain unrestricted access to years of private memories gathered elsewhere.

Security​


And security becomes much stranger.

Connected agents already face the problem of indirect prompt injection: malicious instructions can be embedded inside content an agent reads, causing it to confuse untrusted data with legitimate instructions. Microsoft, among others, now documents indirect prompt injection and memory/context poisoning as explicit agent-security threats.

Now place that problem inside a hostile multiplayer world. A sign on a wall could contain text intended for the agent rather than the fictional character. An NPC could attempt to manipulate its instructions. Another player could construct content specifically designed to alter its behaviour. Amalicious game could attempt to contaminate memories that the agent carries into its next environment.

The fantasy treasure chest has somehow become part of the cybersecurity threat model.

Ownership​


Ownership creates similarly awkward questions.

  • Who owns learned strategies?
  • Who owns the memories?
  • Can publishers inspect them?
  • Can tournament organisers verify them?
  • How do you prove which model and tools were actually used?

And perhaps most strangely: how much of an agent can you replace before it becomes a different agent? Change the model? Probably still the same agent. Erase its memories? Less obvious. Copy every memory and skill into another instance. Now we are doing philosophy because apparently video game matchmaking was not complicated enough.

There will not be clean answers at first. That is usually a sign that something interesting is happening.

From Playing Games to Raising Players​


None of this means persistent cross-game agents are inevitable.

Several important pieces remain speculative. There is no widely adopted protocol through which unrelated commercial games accept player-owned agents. There is no portable agent identity recognised across publishers. Portable memory is an emerging interoperability problem, not a solved one.

Current agent arenas are experiments, not evidence of a mature spectator sport.

And sufficiently capable generalist agents may remove the need for special game integration by simply operating the same interfaces humans already use.

But enough pieces now exist to make the larger question worth asking. Game-playing agents exist. Multiworld agents exist. Persistent skill libraries exist. Game-based agent benchmarks exist. Competitive agent arenas exist.

Protocols can separate intelligent agents from the applications they operate. Work on portable memory is beginning. The extrapolation is not that any one of these projects automatically becomes the future. The extrapolation begins when those ideas converge.

For decades, games have contained their artificial players. Perhaps some future games will instead become places that artificial players visit. Your agent could arrive carrying memories from worlds its current opponent has never seen. It could develop strategies you never explicitly programmed. It could build a competitive history across games made by unrelated developers. Its underlying model could change without necessarily erasing its identity.

And your own role could gradually shift from controlling every action to shaping the intelligence making those actions.

At that point, asking someone which model they use may become less interesting than asking the question we currently reserve for players:

What has your agent learned?
 

Thread statistics

Created
Llorenc Ballester,
Replies
0
Views
3
Back
Top