Roadmap

What's proven. What's tested. What's still being built.

Every capability on this page carries an honest label, because a roadmap that flatters itself is worthless to the person relying on it.

Dates on completed work are what happened. Dates on unfinished work are targets, and we say so.

Every milestone, at a glance.

Milestone, date and status in one table. The sections below give the detail and the evidence.

MilestoneDateStatusEvidence
Bounded understanding of a real codebaseCompleted — Q1 2026ProvenExercised on 13 real, widely-used projects, including one with 4,342 source files.
Verified self-repair with clean rollbackCompleted — Q1 2026Proven109 verified repairs, 0 observed cases of anything being made worse, 100% clean restore when uncertain.
Catching silent, non-crashing mistakesCompleted — Q1 2026Proven10 of 10 planted silent faults caught and repaired; every fake repair rejected.
Refusal and escalation instead of guessingCompleted — Q1 2026ProvenEvery recorded outcome lands in one of two columns: repaired and confirmed, or safely referred.
Impact analysis before any changeNow — Q2 2026TestedBehaves reliably in controlled runs; being measured across more real projects before it is called proven.
Learning that carries between sessionsNow — Q2 2026TestedRetention and reuse verified in controlled runs; long-horizon behaviour still being measured.
Early access with your own bounded authorityTarget — Q3 2026In development
Focus Domain Intelligence for your workTarget — Q3 2026In development
Continuous watch over live systemsTarget — Q4 2026In development
Shared understanding across a teamTarget — Q4 2026In development
Cross-domain reasoning beyond softwareNo date yetResearch direction
Long-horizon autonomyNo date yetResearch direction

What the labels mean

Four honest states.

Proven

Measured on real, unfamiliar software with results published on the Evidence page.

Tested

Works in controlled runs with known answers; not yet measured across enough real projects to publish as proven.

In development

Being built now. Behaviour exists in part; dates are targets, not promises.

Research direction

Deliberate direction with open questions. No date, and we will say so plainly until there is one.

Proven

Measured, published, repeatable.

Measured on real, unfamiliar software with results published on the Evidence page.

  1. Completed — Q1 2026

    Bounded understanding of a real codebase

    Grace reads a project it has never seen, builds persistent context across files, sources, relationships and history, and holds that understanding between sessions.

    Exercised on 13 real, widely-used projects, including one with 4,342 source files.

  2. Completed — Q1 2026

    Verified self-repair with clean rollback

    Faults are diagnosed, repaired, and checked independently of the process that produced the repair. When a fix does not hold, everything is restored exactly as it was.

    109 verified repairs, 0 observed cases of anything being made worse, 100% clean restore when uncertain.

  3. Completed — Q1 2026

    Catching silent, non-crashing mistakes

    The dangerous bugs are the quiet ones: the software runs, it just returns the wrong answer. Grace compares behaviour against the last version known to be correct.

    10 of 10 planted silent faults caught and repaired; every fake repair rejected.

  4. Completed — Q1 2026

    Refusal and escalation instead of guessing

    When evidence is insufficient, Grace hands the problem to a person with what it found, rather than attempting a change it cannot justify.

    Every recorded outcome lands in one of two columns: repaired and confirmed, or safely referred.

Tested

Works in controlled runs. Not yet published as proven.

Works in controlled runs with known answers; not yet measured across enough real projects to publish as proven.

  1. Now — Q2 2026

    Impact analysis before any change

    Before acting, Grace traces what a change touches — the chain of consequences, the blast radius, and the areas it cannot see — and shows that map to you.

    Behaves reliably in controlled runs; being measured across more real projects before it is called proven.

  2. Now — Q2 2026

    Learning that carries between sessions

    Outcomes, successful approaches, failures and open unknowns are preserved, so the same ground is not re-learned every time.

    Retention and reuse verified in controlled runs; long-horizon behaviour still being measured.

In development

Being built now, with target dates.

Being built now. Behaviour exists in part; dates are targets, not promises.

  1. Target — Q3 2026

    Early access with your own bounded authority

    You set exactly what Grace may touch and what always needs your approval, and every consequential action leaves a receipt you can inspect.

  2. Target — Q3 2026

    Focus Domain Intelligence for your work

    Grace narrows to the domain you actually operate in — your systems, your vocabulary, your constraints — instead of behaving like a general assistant.

  3. Target — Q4 2026

    Continuous watch over live systems

    Moving from repair on request to noticing a problem forming, verifying it is real, and either fixing it inside its bounds or raising it with evidence.

  4. Target — Q4 2026

    Shared understanding across a team

    One accumulated understanding of your systems that several people can work against, with each person's authority kept separate.

Research direction

Honest about what we don't know yet.

Deliberate direction with open questions. No date, and we will say so plainly until there is one.

  1. No date yet

    Cross-domain reasoning beyond software

    Applying the same evidence-bounded loop — understand, act inside limits, verify, learn — to work that is not code. Open questions remain about how verification works when there is no test to run.

  2. No date yet

    Long-horizon autonomy

    Work that runs over weeks rather than minutes, with authority that stays bounded the whole way. We will not put a date on this until the safety case is settled.

How this page is kept

Labels move in one direction only when the evidence says so.

  • Nothing becomes "proven" until it has been measured on real software Grace had never seen, with the result published.
  • A missed target date stays visible and gets re-dated rather than quietly deleted.
  • Research directions carry no date at all, rather than a comfortable guess.

Early access

Work on something measurable with Grace.

Request early access