Agents write great code and terrible stories. This workflow makes them show their work: specs before code, proof after every step, and a review you can actually trust.
You get fast code and unreliable stories about what happened. These three failures show up on every real project:
“It wrote that yesterday. It doesn’t remember anything today.”
Every new session starts from zero, so the agent re-asks what the repo already answers and re-builds what already exists. The fix: state lives in files, not sessions.
“It said it tested it. I can’t tell, and I’m about to trust it in prod.”
“I tested it” is a story, not proof. The tree can move after the review, and evidence can describe content that no longer exists. The fix: evidence is re-run, never recounted.
“The review depth depends on the agent’s mood more than the change.”
Big migrations get skimmed, small fixes get three passes, because review depth follows whoever is on shift. The fix: depth comes from the diff, not the mood.
No dashboard to learn, no process to memorize. You talk to Hermes the way you already do, and the agency handles the discipline:
In plain words: “add SSO login to the web app.” An agent turns it into a spec with scenarios anyone can read and argue with.
a spec, not a guessNo code is written until the change is validated. The checks are plain commands you can run yourself, and the diff is reviewed before the next stage starts.
no code without a validated changeThe builder writes the code and the tests together, the test coming first by rule. A behavior change without a regression test does not ship.
test-first, by hard ruleA reviewer attacks the diff. QA runs the scenarios against the real app. Only they can block, and nothing advances on a self-report.
reviewer + qa, independentlyState resumes from files, not memory. You can close the laptop and pick up exactly where things stood.
Every claim binds to the tree it validated. If the tree moved, the evidence bounces and the change gets re-reviewed.
Review approved, QA passed, evidence recorded. Nothing less. You never argue about what “done” means again.
The autonomous run specs, plans, builds, reviews and QA’s the change, then ends in a PR for the morning. Assumptions get flagged, not hidden.
No. One command installs it, and you talk to Hermes in the same plain words you already use. The process adds the discipline; it doesn’t add a tool to learn.
There are three task levels. A typo or a rename skips the loop entirely; a bounded bug fix takes a fast path with a regression test; features and contracts run the full loop. Strictness lands where it protects.
No. The agent personas come with the install, and they take work off you rather than asking you to learn their habits.
It needs openspec ≥ 1.13.0 on your machine, and you have to read the receipts it produces. That’s the whole trade.
Your next feature can run through the loop tonight. You’ll wake up to a PR, not a fire.
bash <(curl -fsSL https://raw.githubusercontent.com/mcabreradev/hermes-sdd-agency/main/install.sh)