AI Agent Systems: Memory and Evaluation

Follow state, attribution and evaluation through small agent systems.

Start with state and failures

BotSim introduces a small local community. The orchestration article follows streamed events, failures and cancellation. Both distinguish deterministic test providers from live language models.

Test coordination

The swarm experiment measures neighbour coupling in a tiny grammar. Mixture of Perspectives separates hard constraints from preference disagreements. Neither fixture establishes general intelligence or better decisions in the world.

Evaluate revision and training

Read self correction before adding repeated answer revisions, then self rewarding training for preference loss and independent evaluation. The Minecraft article is a personal game building account, not a model benchmark.

Articles in this reading path

All reading paths