Contribution guide
Make the result reproducible.
External results are not published yet. These are the requirements for a comparable submission.
Include the whole run.
- Record the source revision, world and scoring versions, provider model ID, prompt and adapter configuration.
- Use five fresh worlds per agent and episode with identical judge models, seed, temperature, step, token and time budgets.
- Include every attempted run, failure and incomplete result, alongside the sweep manifest, scores and evidence artifacts. Keep failed attempts visible.
- Declare exposure to public answer keys and any agent tuning. Public development cases are not a hidden test set.
- Keep critical failures visible. They exclude a result from ranking. Single runs and external sessions with unverified usage remain experiments.
- Report human interventions and minutes only when directly measured; leave unavailable measurements blank.
Use only fictional data. Never submit client records, practitioner identities, contact details or credentials.
Prepare a submission locally
The source includes docs/LEADERBOARD_METHODOLOGY.md, CONTRIBUTING.md and the sweep command. Automated uploads are not enabled in this alpha. Open a repository issue with your configuration and a link to the complete fictional artifact bundle for manual review.