Pick an outcome.
The loop improves it.
An AI agent tries each action, keeps what improves the outcome, and drops the rest.
- CAC
- Signups
- Landing-page conversion
- Retention
macOS · Apple silicon and Intel · Runs on your ChatGPT sign-in or Gemini
Your AI does the work.
It never learns what works.
Your agents can ship a new onboarding flow, a new price, a new headline before lunch. Whether the number moved, and which change moved it, ends up in your head, if it's anywhere. The agent never gets it back, so the strategy for moving the number stays yours to pick.
So you accumulate output. Never knowledge.
You point it at a number. It does the rest.
You do two things: pick the number, and set the limits it must not break, no drop in revenue, no rise in support tickets. After that it runs on its own, and comes back to you only when an action can't be undone.
Set the target
Pick the number you want to move, and the lines it can't cross.
Let it run
It tries an action, watches the number, keeps what works, drops what doesn't.
See what's true
It settles the prediction: what moved the number, and what did nothing.
Pick the next move
The next action comes from what held up. You can override or stop at any time.
it keeps turning on its own, from step 02.
The agent does the work. The harness keeps score.
Point Nimrobo at one number. Every run, it reads back what its earlier runs proved and disproved about that number, decides what's worth trying next, and states what it expects before it acts. The real result settles that prediction and joins the record the next run reads from.
- north star + trend
- guardrails
- open experiments
An agent on your Mac, on a leash you hold.
The agent runs shell commands on your Mac. Three separate settings decide how far it can go on its own: which files and networks a command can reach, how often it stops to ask you, and how fast you can kill it. You set all three.
What can any one command actually reach?
By default, almost nothing outside the folder you point it at. Reaching further is a ladder: every rung up costs an approval.
Outcome folder + /tmp · network blocked · secrets sealed
Writes anywhere · network open · secrets still sealed
Full machine access · your approval on every call
Not a promise the model makes: a macOS Seatbelt profile the kernel enforces.
Sandboxed commands run inside Seatbelt, Apple's own OS sandbox. A blocked write fails at the operating system, not because the agent chose to behave. A project config can widen a run's room but can't un-protect your secrets, and sudo, rm -rf /, and sandbox escapes are blocked outright.
How closely do I have to watch it?
A separate setting from the sandbox: one you move, from fewest-approvals to approve-everything to read-only.
Switch mid-run and it reaches subagents already running, not just the next one.
Can I get in its way while it's running?
At any second. It's never a black box you have to wait out.
Steer
Type while it's running. It picks you up at the end of the current step, not after the whole turn finishes.
It asks you
When a call is yours, it stops and asks. Dismiss it and the turn halts.
Stop
One button aborts the run and kills the process group.
Start with the number you own.
An outcome loop needs a number you can move and can measure. Four templates ship ready to run, each one naming the metric, the system it is read from, and the line it will not cross to move it.
Every plan ships the whole harness.
Nothing in the loop is held back. All that scales is how many outcomes run at once, whether we cover the model, and how much premium web search you get.
One outcome, on your own model.
- 1 outcome at a time
- Your own ChatGPT subscription
- No premium web search
Unlimited outcomes, on your own model.
- Unlimited outcomes
- Your own ChatGPT subscription
- 250 premium web searches / mo
Unlimited outcomes, plus model credits.
- Unlimited outcomes
- $15 of AI credits for Gemini models
- 500 premium web searches / mo
In every plan, including Free
- The whole harness: north star, levers, hypotheses, experiments, verdicts
- Default / Manual / Plan modes with per-action approvals
- Skills, subagents, and unlimited runs
No checkout: install the app, then upgrade inside it when you need a second outcome. Compare plans
The fine print, without the fine print.
Only capacity. Nothing in the harness is gated. Free is $0/mo with 1 active outcome. Starter is $5/mo with unlimited outcomes and 250 hosted web searches a month. Plus is $20/mo with unlimited outcomes, 500 searches, and $15 of Gemini credits. Create a project past your limit and it’s kept but locked: nothing is deleted, and it unlocks the moment you upgrade or drop back under.
Nimrobo is macOS-only. The shell sandbox that confines what a run can touch is built on macOS Seatbelt, and we won’t hand an agent an unconfined shell.
Yes. Sign in with a ChatGPT Plus or Pro account and you don’t need an API key.
Only if you let it. Every run has a mode (Default, Manual, or Plan) and any action that reaches past the sandbox stops at an approval: Deny, Allow once, Allow this run, or Always allow. Escalations ask every single time, with no permanent allow.
Your outcome folders, runs, sessions, artifacts, reward snapshots, and provider credentials stay on your Mac by default. Model prompts go to the provider you chose; account and plan checks and the skills catalog reach our servers. Git push happens only if you add a remote yourself.
No. A coding agent does the work inside Nimrobo; Nimrobo is the structure around it: the number you’re trying to move, the record of what past runs proved and disproved, and the verdict on every prediction. A coding agent finishes a task and stops. Nimrobo checks whether the task moved the number, then uses the answer to choose the next one.
An outcome project is one number and everything the agent has learned about moving it, kept in a git-tracked folder on your Mac. A lever is a kind of change that might move it: pricing, onboarding copy, email timing. A hypothesis is a prediction registered before a run: this change, this much, this soon, for this reason. An experiment tests one hypothesis and returns a verdict: supported, refuted, or inconclusive. Every verdict stays on the record.
It’s a local-first, single-machine app with no built-in collaboration. Outcome loops are just git-tracked folders, so pushing one to a shared remote is how people share them.
Start your first outcome.
Pick one number. Nimrobo works on it every day and builds the record of what actually moves it.
macOS · Apple silicon and Intel · No credit card