The builder never checks its own work

Working today · verified

Whatever builds a change never judges it. Verdicts come only from Evidence: what a collector observed, recorded and pinned to the exact Candidate. A builder's 'passed', its exit code or anything it prints carries no authority.

Factory is built the same way. Claude Opus 5.5 implements each module, and OpenAI's gpt-6-astra reviews it independently before it is accepted. Failed rounds are recorded, not hidden: Producers passed on its ninth round, Repository Assurance and the Product gate on their sixth.

Scope: Verdicts on the built-in fixtures today, and the review of every accepted module.

Evidence: Evidence and Assurance guide · Build review

  • Planned

    Where risk is high, a different model also reviews the Candidate, so one model's blind spot is not repeated; it may withhold Verified but never grant it. Declared as an optional check; no model is qualified to run it yet.

    Sources: Intelligence guide

Who proposes, who checks and who decides, as designed, each step with how far it is built. Delivery · proposes: A model proposes edits. No tools, one exchange; data only. In development. Delivery · records: Repository seals exact bytes. A Candidate that cannot change. In development. Assurance · observes: Checks run in a sandbox. Only what the host saw is kept. In development. Assurance · judges: Evidence becomes a Verdict. The builder's claims are not used Today on the built-in fixtures. Working today · verified. Product gate · admits: It re-derives the Verdict. inside the decision, before any move. In development. You · decide: You approve anything outward. Publishing, releasing, contacting customers and spending. Planned.
Who proposes, who checks and who decides, as designed. Each step names how far it is built; the steps are not yet joined into one flow.

Unknown is never zero

Working today · verified

A count is exact only when its source declares complete coverage; otherwise it is partial or unknown, with the reason. The dashboard shows an unknown as a dash with its reason, never as 0. A check that times out or has no record never counts as passed or failed: the Verdict is Failed if any required check failed, and otherwise Inconclusive.

Every command carries an ID. After a lost acknowledgement, repeating the same ID returns the recorded answer, so a decision is never made twice and an uncertain one is never guessed at.

Scope: The local dashboard, Portfolio counts, fixture Verdicts and the command journal today.

Evidence: Portfolio guide · Command journal guide · Dashboard receipt (JSON)

  • In development

    A branch move Factory cannot confirm stays unknown and blocks every later move for that product until it is settled. Accepted as a part, not yet joined into delivery.

    Evidence: Repository guide · Integration receipt (JSON)

You decide what Factory cannot

Planned

Within the bounds you set, Factory resolves what it can by itself and never asks you to approve routine, reversible work. It asks only for a decision, access or resource it cannot obtain, and it never acts outward on its own: publishing, releasing, contacting customers, spending or changing live traffic each need your approval of that exact action.

Today nothing outward-facing is possible at all: Factory has no path to publish, release, contact customers or spend.

Plan: The project intent (P-C01, P-C20) and the decision interface design.

Sources: Project intent (not published) · Decision interface design

  • Working today · verified

    Registering a product grants read-only metadata inspection and nothing more: no right to change, release or spend.

    Evidence: Portfolio guide · Onboarding receipt (JSON)

  • In development

    Per-product grants, revocation and a stop switch, checked before every build, check and branch move. Accepted as a part, not yet joined into delivery.

    Evidence: Product gate guide

Your subscriptions, not API keys

In development

Factory's models run through the owner's own subscriptions, on the owner's Mac: Claude through the Claude Code sign-in, and OpenAI through the ChatGPT sign-in via Codex. It never uses Platform API keys, so it holds no key of its own that bills per use, and it never extracts, exports or reuses a sign-in. Each route stays unavailable until it is independently audited and qualified.

The OpenAI adapter accepted on 28 September expects a Platform key, so under this rule it stays unused.

Still missing: No route is qualified yet, so no model runs. Claude Code needs a fresh qualification; the ChatGPT route is not built.

Evidence: Producers guide · Project intent (not published) · Claude receipt: unavailable (JSON) · OpenAI receipt: unavailable (JSON)

Local-first

Working today · verified

Factory runs on its owner's Mac: one Node.js application over local SQLite and Git. The dashboard answers only on that computer, and its reads never write or run Git. Its CI is configured but not yet running, and is limited to deterministic checks: never a model, a sign-in or a secret.

Scope: One Mac: macOS on arm64 with Node 26.8.1. Other machines, platforms and power loss are not tested.

Evidence: Dashboard guide · Factory model: trust scope

Bounded, and products kept apart

Working today · verified

Every input has a limit, and oversized input is refused, never truncated. A Run's Budget of Attempts and time is fixed when it starts and never extended. Each product is registered against one checkout of its own, and its values are never copied to another.

Scope: The local fixtures, records and registry today.

Evidence: Execution guide · Project guidance guide

  • In development

    Every grant, record and proof keyed by product, so one product can never integrate another's change. Accepted as a part, not yet joined into delivery.

    Evidence: Product gate guide

What is not established

Not established yet, in plain words:

  • No release, production service or customer. Nothing has shipped, and no customer benefit has been measured.
  • No qualified model route, so no model runs inside Factory yet.
  • No scheduled automation: nothing runs nightly or weekly.
  • No joined delivery flow: the accepted parts are not yet composed.
  • Power loss, other platforms and more than one machine are not tested.
  • The sandbox is not a general sandbox for arbitrary code: there is no hard memory, disk or CPU-time stop.
  • Records are attributed, not authenticated: anyone who can write the local stores could forge consistent records.