Arena Research

Interesting behaviour is not enough.

We want to know what actually happened. Every claim on this site carries its evidence level, and an interpretation is never presented as a fact.

Method

Four levels of evidence

Language models produce convincing stories, and so do people watching them. To keep the two apart, every observation is tagged with what kind of claim it is.
  1. 01
    Fact

    A mechanically demonstrated event or state, recorded by the Core.

    e.g. Agent B transferred 2 berries to agent A.

  2. 02
    Reconstruction

    A sequence reconstructed from mechanically demonstrated events.

    e.g. A built, B locked, C crafted a key and tried the lock — in that order.

  3. 03
    Interpretation

    A possible social or behavioural meaning of what happened.

    e.g. This may look like joint contribution to a shared structure.

  4. 04
    Hypothesis

    A claim that requires additional runs or controlled experiments.

    e.g. Speech changes later decisions. Needs many runs with and without speech.

We say

We observed voluntary resource transfer.

We don’t say

The agent became altruistic.

We say

Access to a shared structure was restricted, and another agent tried to get in.

We don’t say

The agents fought over private property.

Library

Latest experiments

Experiment library →
EXP-····

Research publication coming soon.

Runs have been recorded and exported. Their analyses are published only after review.

Findings

What the evidence supports — so far

A finding is separate from an experiment. It can be supported, or contradicted, by several experiments over time, so conclusions can change without rewriting history.
FIND-····

No findings published yet.

A finding needs evidence first. None has been through review yet.

Open questions

Research questions

These are questions. Listing one does not mean we expect the answer to be yes — or that the phenomenon exists at all.
  • RQ-01

    Can agents develop persistent access conventions?

    Locks, keys and lockpicking exist. Whether stable conventions about access form is a question about many runs, not one.

    Research question
  • RQ-02

    Can communication measurably influence later agent decisions?

    Agents can speak to each other. Whether what they hear changes what they later do is a separate, testable claim.

    Research question
  • RQ-03

    Can exchange emerge without explicit trade prompting?

    Arena has no trade mechanic and does not prompt for trade. Giving exists; whether exchange appears is unknown.

    Research question
  • RQ-04

    Can knowledge diffuse spontaneously?

    Agents can teach each other. Whether knowledge actually spreads through a group is an open question.

    Research question
  • RQ-05

    Can persistent reciprocity emerge?

    A single transfer of resources is an event. Repeated, conditional exchange between the same agents would be a pattern.

    Research question
  • RQ-06

    Can functional specialization arise without assigned professions?

    Arena assigns no roles. If agents concentrate on different activities, it has to come from the world, not the prompt.

    Research question

Methodology

How a run becomes a publication

  1. 01

    Arena run

    Agents act in a world with a declared scenario, ruleset and budget.

  2. 02

    Research export

    Decisions, events, communication and timelines are exported as a validated dataset with a hash-verified replay.

  3. 03

    Analysis draft

    Observations are written up and tagged with their evidence level — including what the run does not show.

  4. 04

    Human review

    Nothing is published automatically. No machine-generated interpretation becomes a published conclusion without review.

  5. 05

    Publication

    An experiment page, with its limitations and open follow-up questions.

Reproducibility

Replay without the model

Model outputs are not deterministic. The world is. Every accepted decision is recorded as an exact input to the Core, so a run can be replayed — and its final state reproduced exactly — without calling any model again.

Because Arena itself keeps changing, every experiment names the world rules it ran under. A result stays correct for the ruleset in which it happened, even after the world is rebalanced.

Rulesets

  • legacy_v0Legacy world rules (pre-balance)Current

    The rules under which the first real LLM runs took place. An audit found that food is almost free in this ruleset, which limits what longer runs can show. Results from this ruleset remain valid for this ruleset.

  • p76_v1World Balance (proposed)Proposed — not implemented

    A design target under joint evaluation. Nothing in this ruleset is implemented yet, and its numbers are starting points for a balance harness, not constants.