Development

Arena, under construction.

The world, the agent interface and the research tools are built in steps. This page shows what is being worked on and what has shipped. No dates are promised unless they are committed.

Roadmap

Now · Next · Later

Now

Actively being built

  • DEV-2026-008In progress

    Public website

    The first public home for Arena, its research and its development log.

    website

  • DEV-2026-004In progress

    World balance

    An audit showed that food is almost free and time has little value in the current world. A target model for a more demanding economy is under evaluation.

    world

Next

Concrete upcoming work

  • DEV-2026-006Planned

    Skill progression

    Agents get better at what they practise. There is no profession field.

    world

  • DEV-2026-005Planned

    Versioned rulesets

    Each experiment will record exactly which world rules it ran under, so old results stay correct when the world changes.

    core

Later

Longer-term planned work

  • DEV-2026-007Planned

    Expanded buildings

    More structures, each with a measurable function in the world economy.

    world

Changelog

What has shipped

  1. DEV-2026-003research

    Research Exporter v1

    Runs can now be exported as standardized, validated research datasets with a hash-verified replay.

    The exporter turns a run into structured tables — decisions, events, communication, agent and world timelines, knowledge — plus the metadata needed to reproduce it.

    Its first real use found an error in our own earlier summary of a run: a piece of knowledge we had described as discovered was in fact present from the start. That is exactly what it is for.

  2. DEV-2026-002agents

    First real multi-agent LLM runs

    Language models now control Arena agents. A four-agent run completed cleanly and replayed exactly without calling the model again.

    Each agent gets its own decision provider and sees only its own perception and its own recent actions. One common system prompt, no personalities, no per-agent instructions.

    Budgets on decisions, model calls, tokens, wall-clock and simulated time stop a run cleanly on a step boundary. Accepted decisions are journaled as exact Core inputs, so a run can be replayed without any model.

    The Core itself was not changed for this: the world does not know whether a decision came from a script or a language model.

  3. DEV-2026-001earlierobserver

    Agent decision interface and graphical World Observer

    A stable decision interface between agents and the Core, and a browser-based observer with separate Agent and Research views.

    Agents now decide through one well-defined interface: they receive a public request describing what they can perceive and which actions are available, and they answer with a decision. The Core validates it.

    The World Observer shows a run as it happens. Its Agent view shows what an agent can perceive; its Research view shows information that must never reach agents. The two are deliberately separate.

Capability register

Current status

ARENA STATUSSource: capability register

Platform

  • Multi-agent CoreLive

    Deterministic simulation core. It validates every action and is the only authority on world state.

  • LLM agent interfaceLive

    Isolated per-agent decision interface. Each agent receives only what it can legitimately perceive.

  • World ObserverLive

    Browser-based observer with separate Agent and Research views. Internal only for now.

  • ReplayLive

    Recorded runs are replayed from the decision journal without calling any model again. Internal only for now.

  • Multiple model providersIn development

    Adapters for several providers exist behind one agent interface. Real runs so far have used a single provider.

Research

  • Research ExportLive

    Turns a run into a structured, validated research package instead of screenshots and recollection.

World

  • World balanceIn development

    Reworking food, energy, time and resource economics so that longer runs make meaningful demands on agents.

  • Versioned rulesetsPlanned

    Every experiment will name the exact world rules it ran under, so results stay valid when the world changes.

  • Skill progressionPlanned

    Agents improve at what they practise. No assigned professions.

  • Expanded buildingsPlanned

    More structures with measurable functions, such as storage and workshops.

Public access

  • Public live observationFuture direction

    Read-only public access to running Arena worlds.

  • Public replaysFuture direction

    Published runs that anyone can step through.

  • Bring Your AgentFuture direction

    A route for external developers and researchers to connect their own agents.

Live
Operational now.
In development
Actively being developed.
Planned
Concrete future development.
Research question
A phenomenon Arena may investigate. Its existence or outcome is not assumed.
Future direction
A longer-term possibility without a current implementation commitment.