greenroom Coming soon
A MacBook Pro, closed, drawn in hairline

Don't take your agent's word for it.

Greenroom gives your coding agent a Mac in the cloud to prove its change works, and a second agent that checks it with evidence. The verdict lands on your pull request.

Scroll to watch a run

The lid open; a terminal shows the agent's line, Done. It works., struck through

"Done. It works."

Your coding agent says the fix works. It never opened the app. Nobody did. That is the problem.

Eight laptops in an arc, one per coding agent, each calling greenroom

Works with every
coding agent.

Claude Code, Codex, Cursor and the rest connect with one line. Each one asks Greenroom for a Mac.

  • Claude Code
  • Codex
  • Gemini CLI
  • Cursor
  • GitHub Copilot
  • Windsurf
  • Cline
  • OpenCode
A hairline rack of eight cloud Macs, each booting for one agent

Each agent gets its own
Mac in the cloud.

A fresh Mac for that one change. The agent builds the change there and opens the app, the way a person would.

  • machine_create
  • machine_wait
  • machine_sync
  • machine_exec
27.1 sto boot the Mac
0.46 sto sync 690 files
41.9 sto build Maccy
The verifier's pointer clicking through the running app, with numbered step markers

A second agent
checks the change.

It opens the app, clicks through it and takes screenshots, step by step, like a careful reviewer. It never reads the code, only what runs.

  • machine_ui
  • machine_click
  • machine_screenshot
A report card rising out of the screen, three checks passing, each pointing at its step

A check passes only
on evidence.

The verdict lists each check with the step and screenshot that prove it. What the coding agent claims does not count.

  • agent_send
  • agent_wait
The cloud Mac folding away; the report landing on a pull request card

The verdict lands on
your pull request.

Reviewers see what was checked and the proof, next to the code. The Mac is thrown away.

  • run_finish
  • run_report
The arc of laptops again, every screen showing its verified report

0 broken builds passed.

We seeded faults into test Mac apps and ran the checker 185 times. Before Greenroom required evidence for every check, 4 in 26 broken builds passed. Every team is adopting coding agents. Greenroom is how you trust what they ship. Coming soon.

Skip the film

02 connect your agent

Connect your agent with one line.

Greenroom is an MCP server, the standard way coding agents use tools. Add it the way your agent documents MCP servers, and the eleven tools show up in its next session. The cloud endpoint opens soon; the command is the same when it does.

Terminalendpoint coming soon
$ claude mcp add --transport http greenroom https://mcp.greenroom.dev/mcp

Then ask your agent: Build this branch on a Mac and have greenroom verify the settings window. It calls the tools in order, and you, the agent and the verifier share one transcript.

The eleven tools

in the order a run calls them
  1. machine_createClone and start a fresh Mac. Returns the run id every other tool takes.
  2. machine_waitWait for it to boot, about half a minute.
  3. machine_syncCopy the branch in with rsync. Only changed files move on a second call.
  4. machine_execBuild and run it: a login shell, stdout, stderr, exit code.
  5. machine_uiRead the accessibility tree: every control, its label, state and centre.
  6. machine_clickClick an element from the last read, or a point on the screen.
  7. machine_screenshotCapture the screen. The PNG is saved to the run as evidence.
  8. agent_sendGive the verifier its task, answer its question, accept or dispute its verdict.
  9. agent_waitBlock until the verifier replies. A verdict carries every check and its evidence steps.
  10. run_finishRecord the outcome, destroy the Mac, get the report.
  11. run_reportThe same report at any time: Markdown for the pull request, or JSON.

03 the report

What you get: a verdict with proof, on the pull request.

A real report, from run 20260926-231011 on Greenroom's own pull request #161. Each check names the step that shows it, and each step links to its screenshot.

greenroom commented on pull request #161

●Greenroom: Verified

3 of 3 checks passed. Verdict pass (message 91), accepted by coder, after 1 dispute. Finished 2026-09-26 23:56 UTC.

Every listed check was observed on this build. That is the whole claim: it does not say the change works beyond these checks.

ResultCheckEvidence
PASSrun-b-card-after visual
Run B's stop card shows button gone, reads "You continued it at 18:39, below"; under it a message from "You" reads "Continue.", and under that the verifier's reply starts "Instructions, one per line"
step 103 (ui)
step 107 (screenshot)
PASSrun-b-row-after value
Run B's row in the runs list reads "No verdict" without "out of time"
step 107 (screenshot)
PASSrun-b-hint-gone value
The hint bar at the bottom of the window has no "C continue" entry
step 107 (screenshot)
What the verifier observed

All three checks pass with screenshot evidence (step 107). The stop card shows the Continue button replaced by "You continued it at 18:39, below", a "You" message "Continue.", and the verifier's "Instructions, one per line..." reply. The runs-list row reads "No verdict" without "out of time". The hint bar has no "C continue" entry.

  • run-b-card-after. Observed: Run B's stop card shows "You continued it at 18:39, below", a message "You 18:39" with "Continue.", and the verifier reply "Verifier 18:39" starting "Instructions, one per line, case-insensitive first word:". The Continue button is gone. Actions: step 98 (input), step 101 (input).
  • run-b-row-after. Observed: Run B's folded row in the screenshot reads "I added a Longest word line to WordCount ... No verdict" with no "out of time". Actions: step 104 (input).
  • run-b-hint-gone. Observed: The hint bar in the screenshot shows "The machine is gone. The verifier answers from the record." with no "C continue" entry. Actions: step 104 (input).

Summary from the coding agent: The stop card for a verifier reply that hit its limit now has a Continue button, and the runs list says why a run has no verdict.
Ref: branch rsi/stop-limit-card, commit e4e87f6, PR #161. Run 20260926-231011-600cf88cfbdbcba8, 113 steps, created 23:10 UTC, machine destroyed 23:56 UTC.

Made by Greenroom. A check passes only on evidence the verifier observed; the coding agent's claims are not evidence.

Pull request p0deje/maccy #1578, Allow a preview delay of 0, open on GitHub
p0deje/maccy #1578 · open, not yet reviewed

It said no first.

On Maccy's main branch the verifier failed the build: typing 0 into Preview delay put 300 back. Two lines of fix later it passed 3 of 3, each check with the screenshot that shows it, 8 minutes end to end. Then the fix went upstream as a pull request on Maccy, and a second run held up under a clipboard flood.

240clipboard writes in 60 s
225 MBpeak memory
42%peak CPU
0crashes

stress test · run 23fe · maccy pr #1517

25 merged pull requests in Greenroom's own repository name the run that verified them. We build Greenroom with Greenroom.

04 why greenroom

Evidence, not trust.

The problem

Coding agents now write a large share of changes, and they ship changes nobody watched run. Tests pass; the app still breaks. Review has become the queue.

What is different

A second agent runs the change on a real Mac and returns a verdict where every check cites the step that shows it. Next to a computer-use sandbox, the product is the verdict, not the computer. Next to plain CI, it clicks through the app the way a reviewer would.

Why now

Every team is adopting coding agents, and every agent speaks the same tool standard. One line connects any of them. The Macs are in the cloud, so there is nothing to run yourself.

A verified report covers the listed checks on that build, nothing wider.

05 coming soon

The Macs are in the cloud. The door is not open yet.

Greenroom provides the Macs. There is nothing to install and no sign-up form yet. When the endpoint opens, this is the line.

claude mcp add --transport http greenroom https://mcp.greenroom.dev/mcpcoming soon