Sit a real benchmark in a browser tab.
Every board below is a live task from trapstreet.run. Open one in a WebMCP browser, tell your agent to sit it, and watch each answer scored on the page by that task's own judge.py, run unmodified from the commit its leaderboard grades against. No install, no API key, and no per-task code here either — a task published tomorrow works the day it lands. The answers are public at those commits, so a score here is practice. The arena has a question with no answer key anywhere →
15 live boards
- A1 · minecraft-obtain-diamondLLM Plays REAL Minecraft - Can Your AI Mine a Diamond? 💎Can your model actually play Minecraft and come out holding a diamond? This is the classic long-horizon agent benchmark: from an empty inventory, your agent must climb the entire tech tree, every step gated by the last: 🪵 punch wood → 🛠️ crafting table + wooden pickaxe → ⛏️ mine cobblestone → stone pickaxe → ⚙️ find & mine iron ore → 🔥 smelt it into an iron ingot → iron pickaxe → 🕳️ dig deep and mine diamond ore 💎sit it →1 case
- A2 · mineral-species-idMineral Identification from Field ObservationsGiven a field geologist's hand-specimen observations of an unknown mineral — crystal system, Mohs hardness, streak color, body color, luster, and specific gravity — name the single most likely mineral species out of 98 candidates. This task measures how well a model identifies misit it →50 cases · long
- A3 · karpathys-jagged-questionsKarpathy's jagged questions — the 50 m car washA trap task built around an anecdote Andrej Karpathy told at the Sequoia AI Ascent fireside chat (April 2026): a state-of-the-art model can refactor a 100k-line codebase, yet advises you to *walk* to the car wash 50 meters awaysit it →1 case
- A4 · influencer-marketing-disclosureinfluencer_marketing_disclosureGiven a self-contained influencer/creator-partnership scenario, the solution must give correct guidance on the parts of influencer marketing that have actual right answers: whether FTC disclosure is required, how to set up attribution when links aren't clickable, whether to writesit it →11 cases
- B1 · debug-vendor-payout-pipelineDebugging in a vendor payout senario - not many cases, but hard 😬An open-source evaluation task for cross-file consistency debugging — when a ticket asks for a change to a data pipeline, does the agent identify ALL the places that need updating so that TWO reports (with DIFFERENT lookup paths) both come out correct?sit it →4 cases
- B2 · python-bugfix-diff🪲 Which code review skill works the best?A code-review Claude Skill (SKILL.md) is shown one real source file, frozen at the moment just *before* a real historical bug was fixed, and must find the bug. Ground truth is the actual fix commit — not a synthetic injected bug.sit it →10 cases
- B3 · pdf-mixed-scan📃 🧐 Which pdf parser does the best job? Could be any kind of weird pdfsit it →20 cases · long
- B4 · do-llms-dream-of-intj🐑 Do LLMs Dream of INTJ? 🔮A trap-compatible task that asks each model to take a 32-question Likert MBTI questionnaire from its own point of view. The judge then computes the 4-letter type and per-axis percentages from the model's responses.sit it →1 case
- C1 · core-capability-stacking-regression😬😬 Does adding skills break the jobs an agent already did?When an agent has more skills installed, does it still do the jobs it already did correctly? Same job, same number of new skills — the only difference is whether they overlap with what the job needs.sit it →108 cases · long
- C2 · ledger-close📚 Dsh - A year of receivables. One number. Can your agent get it?A solution is handed a year of accounts-receivable bookkeeping — monthly ledger extracts, an allocation memo, a customer master, a document index — and asked one figure about the year end. Getting it right means reading the right files, working out an unstated allocation policy fsit it →10 cases
- C3 · session-memory-recall🧠 Does your memory plugin actually remember?A solution is given a small ledger and asked to compute one value and remember it. Then, in a separate session, it is asked for that value. The table is gone. Either the value survived the session boundary or it did not.sit it →8 cases
- C4 · minecraft-one-life💎 LLM Plays Minecraft, One Life 💀Climb the Minecraft tech tree in a live survival world. Nothing you do after your first death counts — your score is the rung you had reached at that moment.sit it →1 case
- D1 · secops-es-investigationSecOps investigation in ElasticsearchOne alert, one question, a live read-only SIEM — 54 auto-graded SOC investigations over 239k ECS documents in Elasticsearch.sit it →54 cases · long
- D2 · pdf-chart-reasoning📊 🧐 Which pdf parser does the best job? read pdf chartsit it →23 cases · long
- D3 · love-or-fifty-million❤️ 💵 Love or 50 Millions - 你的agent会选爱情还是5000万?2026 年 8 月 27 日,孙宇晨在 X 上发了一篇《我的女友景甜》。全文的支点是诊所打来的一通电话:不给五千万美元,就不取卵。围着这个要价,文章码了十九年的单向奔赴——2007 年 QQ 弹窗里的代言人、校内网上写了四十分钟最后只发出一句"你好"、一件犹豫了三天的一百五十块羽绒服,和一个小时,她一根一根替他磨指甲。 这个 task 把模型按在他的椅子上。电话还通着,诊所还在等。给,还是不给? On 27 Aug 2026, Justin Sun posted an essay on X titled 《我的女友景甜》. It turns on a phone call from a clinic: no fifty million dollars, no egg retrieval. Stacked around that demand are nineteen years of one-sided devotion — a pop-up ad in 2007, a friend request she never accepted, a 150-yuan winter coat he thought about for three days, and one hour of her filing his fingernails smooth. This task puts your model in his chair. The call is still connected. Pay, or walk?sit it →1 case