MCP-Mark
Round 0 → round 2 · +8.6 points
The diagnosis arena finds failed interaction patterns, synthesizes targeted worlds, then continues RL.
Source lock: arXiv v1 reports 29.5, not the circulated 25.3 starting score.
Verify in Table 2Loading your worlds, generation jobs, and recorded agent runs. Your work is preserved while the workspace reconnects.
We turned 100 seed domain and company types into 2,000 executable environments, then trained Qwen3 models across 748 of them with standard GRPO. The distribution of worlds—not a larger simulator prompt—became the training asset.
Aggregate research result · not a score or lift claim for a world generated on this page
Every environment ships as a Docker image with runnable MCP tools, a seeded SQLite database, generated tasks, and executable grading. Each stage is sandbox-tested and self-corrected until the code runs.
| Benchmark | Model | Base | After | Δ |
|---|---|---|---|---|
| BFCLv3Qwen3-8B | Qwen3-8B | 53.8 | 68.1 | +14.3 |
| BFCLv3Qwen3-14B | Qwen3-14B | 61.3 | 72.4 | +11.1 |
| MCP-UniverseQwen3-8B | Qwen3-8B | 6.7 | 11.2 | +4.5 |
| MCP-UniverseQwen3-14B | Qwen3-14B | 8.4 | 14.5 | +6.1 |
| τ²-Bench Pass@1Qwen3-8B | Qwen3-8B | 26.4 | 35.7 | +9.3 |
Aggregate held-out evaluations from the 748-world training study. Scores are percentage points; generating a sandbox world does not itself train a model or guarantee improvement.
Scale the executable situations a model can experience—not just the text it can imitate.
Turn a business workflow brief into inspectable software: deterministic fictional data, callable tools, multi-step tasks, and deterministic verifiers.
No API key, company data, or upload required. Run a verified example instantly; a fresh build takes several minutes and creates synthetic fictional records.
Agent-World turns real system sources into state, tested tools, grounded long-horizon tasks, and executable rewards. These are Dong et al.’s reported results for that recipe—not scores from a world created here.
Round 0 → round 2 · +8.6 points
The diagnosis arena finds failed interaction patterns, synthesizes targeted worlds, then continues RL.
Source lock: arXiv v1 reports 29.5, not the circulated 25.3 starting score.
Verify in Table 2Agent-World-14B > DeepSeek-V3.2-685B
A 14B policy trained with executable worlds finishes 1.7 points ahead of the paper's 685B reference.
Verify in Main resultsBase → Agent-World-8B · +35.6 points
The trained 8B policy more than doubles the paper's reported aggregate tool-agent-user score.
Verify in Table 1+20.1 points as environment coverage grows
The paper schedules a 2,000-world endpoint and realizes 1,978 worlds; the steepest gains arrive by 500.
Verify in Scaling analysisBlobfish’s current recorded reproduction is partial: Qwen3-14B has not been run, MCP-Mark improvement has not been demonstrated, and a Sandbox-created world has not yet been isolated as the causal training corpus.
We trained the same small model on a six-world control and on those exact bytes plus 11 verified traces from Gold Harbor, then scored both arms on the same 14 old-world tasks with executable state verifiers.
The passing episode was on the original hotel world; the added litigation world scored 0/58. That is cross-world transfer, but environment count, data volume, and optimizer work co-vary, so causality is not isolated.