#Changelog
Written per release, on the releases page — each one explains what changed and why, which is more use than a list of commits. The entries below are those notes in full, newest first: what was wrong, what it does now, and why the fix took that shape.
An entry that refers to the previous entry therefore points down the page, not up.
#Unreleased
#A palette of its own, a map whose edges meet its nodes, and a header that fits a phone
The interface wore a CSS framework's defaults: slate grounds, an indigo accent, two radial glows behind every view and frosted glass on every panel, set in Inter at eleven and twelve pixels. It reads as generated because it was. The ground is now a near-neutral ink, the accent a soft blue that carries dark text as a fill, the status colours a step softer, and the type is IBM Plex a size larger throughout. The glows and the blur are gone. The mark had kept the old colours, a white A on a gradient of indigo into sky, on the header, the favicon, the iOS icon and every link card; all four now show one flat plate of the accent with the A cut in the ground colour, and the header's copy, which had drifted to thinner strokes than the tab beside it, uses the generator's numbers.
Every edge on the map was drawn 24 units left of the nodes it joined: layoutNodes placed edges at one x and the renderer drew circles at another, so each line entered its circle off centre. Both now read one columnCentre. Reached nodes are a quiet filled disc in place of a sea of violet with a white dot in each, a next step is the bright ring, a blocked node keeps a legible name instead of fading to three quarters, and edges are grey until a node is in focus. On a phone the map opened at 40%, where a name is four pixels tall; it opens at a readable scale and scrolls, and a simulation scrolls to where it starts.
At 390px the header scrolled sideways and Proposals and Docs sat past the edge with nothing to say they were there. A phone gets two rows: the brand and the actions, then the three views across the full width. The proposal cards were flex items with overflow: hidden, so each was squeezed to fit the list and cut its steps off mid-line; the Docs tabs were squeezed the same way until their labels were cut in half. The header's verified count held a glossary button inside its own button, which React reported on every load.
The copy lost what gave it away: an uppercase eyebrow over the Time & cost headline, an icon before every card title, callout boxes with coloured left edges, Title Case in the demo's proposals, "Record the no", hotkey docs that listed two lenses of three, a Docs page that called the columns "areas of work" when they are eras, and descriptions joined to their hints with a second dash.
#Every view says what it found before it shows the evidence
The interface computed what mattered and then printed it in its quietest type. On the map every reached node was the same saturated disc, the next steps that are its recommendations were thin outlines with a nine-pixel cost, and blocked nodes were faded until they could barely be read. Reached is now the quietest filled state, a next step carries the heaviest ring with its cost in the accent colour, a blocked node is a dashed gap that stays legible, and a failing node is ringed in red. One line over the map states the finding: a capability configured and failing its check, or failing that, the next step that reaches the most, with a button that previews it.
The detail panel called a node green "Reached" and, underneath in muted grey, "Configured, never verified". That gap between configured and working is the product's argument, so it is now the panel's first block: amber for never checked, red for failing, with ambit verify as the one filled button. The header's status takes the same colour, so green means something proved it. The outage sentence is body size, with its count in red. My Setup opened on "28 of 28 entries enabled" over a column reading "Enabled" on every row; it now opens on the entries that provide a failing capability and the servers that are enabled and put nothing on the map, and each row says whether what it provides has been checked. Time & cost leads with the saving in its largest type and the opportunity that pays back soonest, and a capability configured and not working is an alert across the page above the figures.
The demo contradicted itself. Time & cost reported 38 checks passing and "E2E on Edge" failing, over a map on which nothing had ever been checked and no node had that name. npm run demo:generate now runs the fixture's checks through the engine's own runner, with true and false standing in for real commands as the fixture stands in for a real config, and the loop snapshot reads its evidence counts off that tree. Two reached nodes are left unchecked, so the demo can show the state the panel exists to flag.
unlockCascade credited a next step with everything reachable from anywhere. On the demo tree six unrelated next steps each "made 18 more capabilities reachable", the same 18, and the unlock simulation lit all of them. It now counts only what depends on the node: Model Routing reaches six, Embeddings four, and three reach nothing on their own.
The link cards were a brand: the tagline beside a graph of coloured dots, naming no tool a reader uses, and every docs page shared it. The home card now asks what breaks if one MCP server goes down, names the runtimes Ambit reads, and draws one node down and the six that go with it. Each docs page has its own card, generated from a one-line reason to open it that build-docs.ts keeps beside the page's title, and carries its own image, dimensions and alt text.
#The hosted site says what it is, in text a crawler can read
An audit of the hosted site found a single URL with about two hundred words on it, rendered by JavaScript, while everything quotable lived on GitHub. The documentation is now published there as static pages under /ambit/docs/: the guide, the FAQ, the reference, the argument and its theory, the roadmap, security, contributing, the changelog, and a new page on Ambit and Jev. Each has its own title, description and canonical URL, heading anchors that match GitHub's, and system fonts so nothing blocks it. Links between documents stay on the site, images are copied beside the pages, and anything else points at the file on GitHub. A test builds the whole set and fails on any internal link or anchor that does not resolve.
The sitemap is generated from the same list, so it holds the home page and the published pages and nothing else. It used to list ?demo=1 views of a page that names the home page as its canonical, which told a crawler two contradictory things. The landing page carries a static summary with links into the docs for anything that does not run JavaScript, hidden before the app mounts so no reader sees it flash. Its title and description lead with what people search for, its structured data carries the released version, author and repository, and its fonts no longer block the first paint.
The welcome screen moved on phones. Its hero was centred with align-items, and a flex item taller than its container overflows at both ends, so when web fonts arrived and it grew, it grew upward under the reader. Auto margins centre it when it fits and top-align it when it does not. Fonts load with display=optional, so text never re-flows. The hosted demo stopped asking GitHub Pages for /api/health, which logged a 404 on every visit, and the primary button's white text now clears contrast.
#Jev on the map, and kept off the authority path
TypeSafe's Jev, a model that answers typed questions with calibrated probabilities instead of text, reached a large share of agent stacks within days of its release, and every one of its MCP servers landed on the map as an unmapped tool server. The curated tree has two new nodes. Typed Judgment, in Model Access, is reached by Jev through its API, an MCP server, or an open clone serving the same endpoint. Local Typed Judgment, in Sovereignty, is reached by a clone on the machine's own hardware, which is the version that keeps the judged state off a third party's servers. Local Typed Judgment has a declared check: it sends one trivial question to a clone on the loopback ports Kev and OpenJev use, or to TYPESAFE_BASE_URL or JEV_BASE_URL when either names this machine, and passes when judgments come back. It never contacts any other host, and refuses a URL carrying userinfo, which curl would otherwise read as a request to whatever follows the @. Typed Judgment has none, because proving the hosted API answers means spending the person's key on a call they did not type.
ambit goal "<sentence>" --judge closes the gap the roadmap named for goal routing: a goal that names none of the tree's words had no route in. When the vocabulary cannot recommend, the sentence goes to a judgment model on this machine as one Choice over every node, and the answer comes back as a suggestion with its probability, or as the three likeliest when none reaches even odds. It writes nothing. It asks only a URL that parses as loopback with no credentials, because the goal is the person's own words and a prefix test would pass http://127.0.0.1:1@host; it defaults to Kev on 127.0.0.1:8009, and a goal the words already cover opens no socket. It is the fifth command that opens one, and the security invariants list it.
The docs say where Jev stops. It is a capability a runtime may consult before it asks ambit can, and never a substitute for the grant: a published test moved its probability of blocking a destructive command with injected text, and a grant a person set in advance has no such input. The hosted demo and the README examples carry a Jev server, so both show the node reached.
#Authority for a window, a tree you can extend, and a decision that reaches you
A standing grant was permanent confirmation or permanent autonomy, so a person pairing closely with an agent for an afternoon either confirmed every command or granted autonomy for good. ambit authority grant <cap> autonomous --ttl=30m grants it for the window and no longer. Expiry is read at decision time and never written: an expired grant decides nothing, and whatever stood before it decides again, a standing confirm as confirm and nothing at all as a refusal. The first cut had an expired grant fall back to confirmation, which meant a row that had ended kept buying something forever; the test that encoded it is the test that now forbids it. A forbidden grant takes no TTL, for the reason a threshold does not, and --by names who declared the grant, as it does for promote and sandbox, without binding the grant to them.
The curated tree is general software development, and a team whose capabilities have other names had nowhere to put them. .ambit/techtree.json in the working directory, or AMBIT_OVERLAY_TECHTREE, is merged over the curated tree before seeding: a new id is a new node with its contract and prerequisites, an existing id extends the curated node's detection and requirements, and override: true replaces it. An overlay is data like the tree it extends.
An agent that drafts a proposal while the person is away from the machine used to wait. ambit dispatch <id>, or --dispatch on propose and approve, pushes the draft to a Slack, Discord or Telegram webhook, an ntfy topic, or any JSON endpoint named by --to or AMBIT_APPROVAL_WEBHOOK: the goal, the cost, what it unlocks and the commands that decide it; once approved, the signed artifact, so what the phone shows is what apply will verify. The channel is one-way on purpose. Nothing that arrives on a webhook mints an approval; the reply comes back through ambit approve on a machine that holds the key, and a chat is not one. Nothing is sent without a URL, and the four commands that open a socket are listed in the security invariants.
The infrastructure scan reads the local Docker socket when there is one, with a single read-only request, so running containers appear beside the manifest's services on a device:docker node, with their image, state and ports. No socket is a quiet absence; a socket file with a stopped engine behind it is a warning. Four capabilities gained declared checks, each one command that reads and changes nothing: whether MCP servers can launch (npx or uvx), whether there is a client to query data with, whether a delivery CLI is present, and whether the mesh or the engine answers.
#The decision, both ways, and what an agent may do
The README promised approval in one click, and on every machine that had not declared a web actor by hand the click failed: the engine refuses a decision from a person the graph does not know, and nothing had declared the person at the browser. The API declares them on the way in now, once, as the person at this machine's loopback port. Refusal had no route at all, so a no made in the browser vanished and the record the next draft learns from was one-sided; the card has a Turn down button, an optional reason, and a Turned down tab, and the engine records who and why. Steps drafted by the engine, which are shaped {id, name, chosen}, printed as their own JSON on the card; they read as a name and what supplies it.
Authority is per action, and the panel said only the capability's own mode. It lists what the capability may do now, each action with whether it asks: read the output without asking, run a command with. What work asked for and never had, which reached the page as a fragility footnote, heads the queue of what to reach, with how many times it stopped work and whether the same cause recurred, since a capability that stopped work four times this week outranks one that would be neat to have. The ways to acquire a capability were a list; they are compared, cheapest first, cost as length on one scale, privacy as a tag, and the one the record of this person's decisions favours is marked, which is the same choice ambit propose would draft.
Two smaller things. A repository missing a server the global config has offers the entry ready to paste, composed by the endpoint that exists for exactly that; the entry still crosses into the config by the person's own hand. And the briefing an agent is given at connect, the MCP resource ambit://briefing, is a tab in My Setup, so what the agent believes about the machine can be read by the person it believes it about. Reading it applies any authority threshold the evidence now supports, as the resource and ambit briefing do, and does not mark the environment briefed. The detail panel says how long a capability's configuration has gone unchanged, once it is two weeks, which is the glossary's decay with its name on it.
#Every state carries its reason and its next move
An audit of the web app for what a person could decide from it found the screens strong on state and weak on why and so what. Reached, next step, blocked, passing and failing were all drawn; what stood behind each, and what to do about it, mostly lived in the terminal. Most of it was already computed.
The governance half had no web surface. The Time & cost page now has one: reached capabilities by whether they act without asking, ask first, or are forbidden, drawn as one bar in the assurance figure's three inks; the grants the engine says have earned a threshold nobody set, each with how often it was confirmed by hand, what passed, and the command that records a threshold, which stays a command because promotion needs a person; standing budgets as spend against ceiling; and sandboxes. The detail panel states the effective mode on every node that has one.
What to reach next is a ranked list with its reasons, the same three ambit next prints: what has blocked work and how often, what reaching each would also reach, and the setup time, with the basis named, observed or structural. The frontier was a set of outlined circles with a cost. The opportunities table keeps its own question, what would pay back, and is titled that.
A blocked node used to be a faded circle whose panel listed needs without saying which were unmet. The panel now says what it is blocked by, every hop up, and how long closing the gap would take; a third simulation draws the gap on the map in amber. Each era's header carries the setup time left to finish it.
The outage cascade painted everything downstream red, so losing one of two providers read the same as losing the only one. The tree payload carries each node's providers and each edge's kind now, and the simulation draws what stops in red and what only loses a provider in amber; the panel and the banner count both. The runtime, which contributes every entry read from its config, is no longer any entry's single point of failure in ambit status or on the page.
A proposal card showed the goal and the steps. Deciding needs four more things the engine had stored and never shown: what it would save, from the economic case; what it costs, from the steps; whether every step has an inverse, which is what ambit apply will run; what it unlocks, from the stored simulation; and how this person has decided on things like it, from the record of approvals and rejections. All five are on the card, and the demo's cards carry the same, written by hand like the rest of the demo.
Evidence says how many runs it rests on, "check passed, 12 of 12 runs", where it read the same for one run and for forty. Failures the runtime reported are classified by the engine and listed on the panel by cause, a 401 as permission, a missing binary as a tool, so "check failing" says why. My Setup's evidence column carries the same. And the week's movement, gained, emergent, lost and diminished, is a strip under the Time & cost title, where the ledger knows what became reachable without anything new providing it, which no per-component changelog can show.
#One map, one list, one count
The web app had two maps and a list of the same things, and the header counted a mixture of them. The tree was one map. My Setup was a second, the machine's entries drawn as nodes in domain columns with nearly every edge running to the runtime, which is a list drawn as a star. The capability list docked on the left was a third of the window, open by default, listing what whichever map already drew. The header pill read "42 of 60 reached" on a demo whose own welcome page said "15 of 33 on the tree", because it counted the entries in with the nodes; on a live machine the list said 170 and the map drew 39.
The map is the tree, and only the tree. My Setup is the list it always was, full width: one row per entry a config declares, with whether it is on, what the engine's check said, and the nodes on the map it provides, each a link onto the map. Repositories and infrastructure, which were tabs inside the side panel, are tabs here, since they are the other two things a machine has. The docked list is gone; Search, on the header and on /, finds anything by name and opens it where it lives, a node on the map or an entry in My Setup. The store holds one list for all of it: the engine's tree, which already carries the entries and the edges from each to the nodes it proves, merged by id with the config read-out, so an entry keeps the config's facts and gains the engine's evidence, and switching views never goes to the network.
The header counts one population per view. On the map it is three segments, reached, next step and blocked, each of them the same control as its legend key, so the count is also the way to see those nodes alone; the glossary calls next step and blocked the informative states, and they used to be absent from the number that led. In My Setup it is enabled of total, over entries.
Selecting a node used to light the whole connected component, both directions, every hop, so a keystone lit most of the map and needing and enabling read the same. One hop now, and apart: what the node needs in teal, what it enables in indigo, with the rest dimmed; hovering previews the same. The transitive answer is the simulation, and the detail panel states it before the button that draws it: "If this went down, 14 other capabilities would stop working." It used to offer ambit impact to copy into a terminal for a question the page could answer. The panel's one list of neighbours, each with an edge word, is two lists, Needs and Enables, keyed with the map's two colours. ambit verify stays a command, because a check executes.
Rows inside a column used to follow insertion order, so height meant nothing while the Docs said it showed how far up the tree something sat. Each column opens in state order, reached first, and a few barycenter sweeps pull every node toward the mean row of what it connects to; edges are curves that leave and arrive horizontally, so a bundle into one node fans instead of converging through everything between; a long name wraps onto two lines where it was cut at eighteen characters with an ellipsis. The Docs say what height means now, which is nothing on its own.
The Shared credentials lens coloured any node whose id contained github, docker, 1password or credential, a string match dressed as an analysis, on a map whose curated nodes never matched it. It is withdrawn; credentials are ambit credentials, the engine's report. The Attention lens is offered disabled, with the reason in its tooltip, until the ledger has recorded something to colour; it used to grey the map and lay a note over it. My Setup's domain columns were named Foundation, Pipeline, Guard and Fortress, none of which is a word the glossary or the detail panel uses, and Foundation is also the tree's first era; a graph without eras names its columns with the domain's own word.
Two figures were being drawn as measurements of nothing. Time & cost showed "0h saved, $0 a year" on a machine whose ledger held runs but no interventions; a saving is a difference between months, so the figure waits for an intervention or a second month and says so. "One provider away from lost" listed most of a machine, because a runtime contributes every entry read from its config and so is each one's single provider by construction; ambit status and the page exclude the runtime from that report, and ambit impact runtime:<name> still answers for the runtime itself.
Narrow screens open on My Setup unless the link names a view: at phone width the map is texture and the list is not. The hero recorder writes the README's two stills as well as the GIF, so the pictures of the product are made the way the product is.
#The map says less, and the panels say each thing once
An audit of the interface for marks and controls that repeat something already on screen, or explain nothing, found a dozen of them.
On the map, a next-step node wore two rings: the outlined circle the legend keys as Next step, and a second, wider one outside it that filled in proportion to a readiness fraction no surface defined. Above many of them sat an amber tag reading Boost, a word in no legend, glossary or document. Both are gone; the outline and the setup cost beside it stay, which is what the legend and the Docs overlay say a next step looks like. The tooltip no longer opens over the selected node, whose detail panel is already saying everything it would, and its last line now says what a click does: it opens that panel, where the simulation is a button. It used to promise that the click would run the simulation itself.
The zoom controls are one control shorter. The percentage reading is the button that returns to actual size, so the 1:1 button and the divider before it are gone, and the lens buttons dropped the small keycap after each name; the tooltips and the Shortcuts tab still carry the keys. The controls also stopped moving. They were sticky inside the scroller, and centring a node carried them off the left edge of the window; a left inset did not save them, because Chrome measures it from the scroller's padding edge and other engines from its border edge. They are laid over the map now and stay where they are.
The detail panel opened with three boxes before its first sentence: a tinted banner for keystone, a bordered one for the check, and a third around the config switch. Keystone and evidence are lines now. A Notes block that said "Not reached yet" under a header reading "Next step" is gone, and the one fact the header cannot carry, a node with no edges at all, is a sentence. The Details list no longer prints state, next, era, era name and a count of setup seconds, which are the status in the header, the column the node sits in, and the cost the map writes as "15m"; setup time is written the same way here. Beside each neighbour the panel said hard-dep or soft-dep, the one place the data model's words reached a reader; it says required, optional or enables. The two commands are lines with a copy control, not bordered rows.
The Proposals overlay lost its search box. A machine holds a handful of proposals at a time and the three status tabs are the filter; a search over five cards was a control that never earned its row.
The Docs overlay's three reference tabs used eleven class names that no rule defined, so "Click a node" ran straight into its explanation, the line samples for Required and Optional were empty spans, and the shortcut keys were styled inline. The rules exist now, and the rules that did exist, for two tables nothing rendered, are gone, along with the tokens nothing read: a copper spectrum, a type scale and a line-height scale no rule used. Reading the Map explained outlined circles twice, under "Solid vs outlined" and again under "The states", the second time mentioning a halo the map no longer draws; it is one section, in the legend's order. The glossary's entry for maturity pointed at a ring around each node that has not been drawn for two releases.
The Vite dev proxy follows AMBIT_API_PORT, the variable the API server already reads, so the two can move together when 3000 and 3001 are taken.
#One word per thing, defined where you meet it
An audit of the interface for house shorthand found five concepts with four names each. The circle whose prerequisites are met was Next step in the legend, one step away in the header, next in the glossary and the frontier in the README. The ● node was a Combo in the legend, a Possibility in the detail panel and a Tech tree node in the docs. The ◈ node was a Server, a Tool server and an MCP server. The panel on the left was List on its button, Capabilities on its tab and the console in the code. None of it was wrong, and all of it defeated the learning a reader had already done. There is one word for each now, and src/client/vocabulary.test.ts fails the build if a retired synonym reaches a user-facing string.
The glossary was a dictionary behind a button: twelve definitions, in no particular order, disconnected from the screen that raised the question — and its own overlay opened by saying "Nine terms carry all the meaning here". The count is read from the file now, the file is ordered the way a reader meets the terms, and the words carry their definitions with them: a dotted underline in the header and the detail panel opens a card with the same sentence ambit help <term> prints, and the map's legend keys and era headers — SVG, where a popover cannot go — say it in a tooltip. Six terms that were on screen and in no glossary were written: keystone, tool server, attention, proposal, and the two halves of required-versus-optional.
Keystone and bottleneck were one idea under two names counted two ways: the map marks three or more dependants, ambit status counts the combos a capability unlocks. The glossary now states both rules under one term rather than implying they agree, and ambit help matches a concept's whole text, so looking up either word finds it.
Reached was the word for an MCP server that is switched on: the detail panel read "Tool server · Reached" directly above a switch labelled "Enabled", two words for one fact, disagreeing. My Setup says Enabled and Disabled, which is what its statuses mean; the Tech Tree says Reached, Next step and Blocked, which is what the legend beside it says. The two filters that had the same five words a few pixels apart — one over the map, one over the list — now each say which they act on.
#The half of the product you could not get to
An audit of the web app for things that worked and could not be found turned up two kinds of problem: state the URL could carry and nothing could produce, and endpoints and actions with no control anywhere on screen.
Time & cost — where a person's hours went, what they cost, and what would pay back fastest — rendered only when the store held demo data, so the tab did not exist for anyone running Ambit on their own machine. GET /api/loop now composes it out of the reports the CLI already prints (status, attention, opportunities, roi) plus the one query none of them makes: hours in the loop month by month, annotated with the proposal that landed. A ledger with nothing in it says so and names the two telemetry bridges that fill it, rather than drawing a page of zeroes that reads as "your time costs nothing".
Every view was linkable and no link could be made. The address bar now follows the view — graph, selected node, lens, filter — and Share copies it. ?treeFilter= had been readable from a URL and settable from nowhere; it is a control over the map, and compact, which was in the accepted list and in no renderer, is gone: it fell through to "frameworks only", a filter that hid the graph and could only be reached by typing it.
The legend's keys were buttons that spotlight a kind of node, with nothing to say so — they carry the affordance now, and a row that says what clicking does and how to clear it. A node's tooltip says which simulation a click will run, because the unlock simulation lives on faded nodes that read as scenery. The map says when it is live, since it redraws itself whenever the graph is rebuilt elsewhere. The attention lens explains itself when there is nothing to colour.
/api/repos/scan and /api/infrastructure/scan had no caller in the client: the work was done on request and thrown away. Both are tabs in the side panel now. toggleMcpEnabled had no button, so the one edit the browser may make to a real config — enabled: true|false on an entry that already exists — was unreachable; it is a switch on an MCP node when an engine is behind the page. loadFromJSON had no drop target, so "what does this look like for my setup" meant cloning the repository; the welcome screen takes an opencode.json and maps it in the tab, uploading nothing.
Two things were broken rather than hidden. Every "Map" button on the opportunity table matched the row's id against three hardcoded strings and fell through to the same fallback node, so all three led to Wrangler; each row now names the capability it prices, and the button switches to the map with the unlock simulation running. The demo's attention counts were seeded into the store's initial state, so a live machine briefly coloured its map from the fixture.
#The docs say each thing once
The README had two install sections that named Homebrew, Codespaces and the MCP registration twice each, and a nine-bullet reference on delegation records sitting in what is otherwise an introduction; the install is one section now and the record reference moved to the deep dive. The deep dive carried five "real-world walkthrough" scenarios whose console output the engine has never printed — cascade risk: CRITICAL, Available Frontier: 142 capabilities — in a voice that belonged to a brochure; they are gone, and the three honest cases in the README stand. Its canvas section, written in set notation with glowing emerald green, is a paragraph on what the map may and may not touch. The roadmap's §12 and §13 were specifications of things since built, three thousand words of them, and are now the decisions each part made and why, with the deep dive as the reference; its status table stopped saying nothing promotes a capability on evidence and nothing records use, both of which have shipped. The why essay's postscript, which listed what existed as of a version that no longer does, points at the roadmap instead. CONTRIBUTING had its list of good first contributions twice and the security invariants a third time; once each now, with the invariants in the agent guide alone. Roughly five thousand words fewer, and no claim in any of them that another file contradicts.
#The front door shows measurements
The landing page drew four circles labelled LLM, MCP, Tool and Goal for a product whose whole claim is that it measures things. It now draws the two measurements the demo makes, from the demo's own example data and labelled as such: a year of a person's hours with the two acquisitions named beside the drops they caused, and seven eras as small multiples on one scale with the reached fraction beside each bar and the next step drawn hollow. The sparkline moved into a shared figures.tsx so the landing and the dashboard are one drawing hand, and its annotations — which the dashboard had been marking with an unlabelled tick — are now named, staggering onto a second row where two would overprint.
The map's era columns say Era 5 · 0 of 5 under the name with a three-pixel fraction bar, all seven on one scale so a short bar is a small era rather than a poorly filled one. The status pill draws the same fraction it states, where a green dot had stood for nothing. The dashboard's money column gained the scale header the other columns already had, the interruptions column names its unit, payback marks say mo, the assurance caption carries 38 of 49 proved · 78%, and acquisition options put the dollar figure first in a fixed column so two alternatives compare at a glance. The capability list's filled, bordered status pills — three encodings for a two-valued fact, sixty times over — became a dot and a word.
#The map looks like one product
The client had two palettes — the slate-and-indigo tokens, and a cyberpunk cyan/magenta/neon-green the tree painted its nodes with — and fifteen CSS classes that no rule defined, so the zoom controls, the toast, the small buttons, the dashboard's filter tabs and the sidebar toggle all rendered as browser defaults. Every kind of node now has one colour, read from App.css by the map, the detail panel and the docs alike; the missing rules exist; the favicon and the social card are drawn in the same indigo as the app. The landing page renders on its own, without a capability list reading "(0)" behind it. The time-and-cost view sits beside the capability list instead of under it. The lens switcher moved from the top bar onto the map it colours. The map opens fitted to the window, so the seventh era and the runtime column are on screen instead of past the edge; the first-run card sits top right instead of on the legend; the legend describes the view it is under. The proposals panel says "waiting for your approval" and "approve and sign" where it said "AWAITING OPERATOR RATIFICATION" and "Ratify & Sign Policy", and signs as the web surface rather than as the author. Muted text passes AA contrast, every control shows a focus ring, and motion respects the reduced-motion setting.
The README's line is now also the page title and description, the package description, the citation title, the agent guide's first sentence and the social card. Thirteen unreferenced brand PNGs and a byte-identical duplicate of the social preview are gone; npm run assets:generate draws the card and the favicon and nothing else, and the README screenshots are captured with the first-run guide off. ambit status draws its per-domain rows as bars, so the first-run report and the command print the same thing; bootstrap.sh no longer keeps its own copy of that rendering, nor a tt alias. The launch kit moved out of docs/ into .github/.
#The numbers are drawn, not narrated
The demo's Time & cost page said "Today: 43 interventions a month, 8.6h of your time, $2150/mo. After: 0.8h a month, saving $1935/mo. Pays back in 0.6 months" — five figures in a sentence, three times over, with no way to compare one row against the next except by holding them in your head. The figures have not changed. They are now on shared scales.
What to set up next is a table with a graphic in each cell: the hours as a dumbbell from where they would land to where they are, every row on one 0–9h axis drawn once in the header; the money recovered as a bar on one scale; the payback as a mark against the month it has to beat. The old payback bar filled in proportion to nothing at all.
The headline saving is drawn as the thing it is. Twelve months of hours run under it with January held as a dashed reference, and the shaded band between them sums to exactly the 41 hours the number claims. Two months carry the acquisition that caused the step down. The forecast tile shows 37 predicted against 41 spent on one axis, with the labels on separate baselines because the two marks sit closest together precisely when the forecast was good. Four status chips became one part-to-whole bar; green and red never share an edge, because under deuteranopia they are the same colour at ΔE 5.6, so the grey segment separates them and every segment carries a written label.
Interruptions worth removing and decisions worth keeping were two lists side by side, which put a gap through the middle of the comparison the section exists to make. One chart now, one hours axis, sorted, with the keeper in the recessive grey under a hatch that survives greyscale and forced-colors.
The attention lens on the map painted a magnitude in two hues split at an undocumented count of twenty, and its legend never mentioned it existed. It is one hue in four steps now, brighter with more, monotonic in luminance and each step clearing 3:1 against the canvas, with bins taken from the data and printed in the legend under the unit they count. Its demo data pointed at three nodes no fixture contains, so pressing 2 on the view the demo link opens dimmed the map and warmed nothing; the counts now name real capabilities and are the same four the Time & cost page prices, so the two surfaces read one ledger.
#One fact, written down once
An audit for redundancy turned up two defects and a set of places where the same rule had been copied until the copies disagreed.
Recording a spend against a capability with no budget created a budget with a ceiling of zero, so a single recorded cent refused every later spend for ever — through a row ambit budget did not list, because a report of standing budgets reasonably skips the ones with no ceiling. Spending where nothing was delegated is now an absence of a budget and says so.
ambit sync carried human interventions and not the work runs they hold a foreign key to, so every intervention was silently skipped on import and counted as skipped. Runs, capability use and outcomes now travel, which matters more since successful work became the evidence that earns authority: a rebuilt container was losing exactly the record that had been earning it.
The graph-summary query had five copies, two of them byte-identical inside the MCP server, and the visualiser's counted reached as not-locked where every other surface counted unlocked-or-active — agreeing only because a third state has never been added. There were four wordings of "the graph is not seeded" offering three different fixes, two identical intervention-kind vocabularies under different names, a third partial copy of the same list, and five copies of the default config mapping differing only in the two words before "{type} server". All of it now comes from src/engine/vocabulary.ts and one parameterised mapping builder.
The formatter dropped nested objects entirely, so ambit can printed six of its nine fields and the governing grant and per-target evidence were visible only with --json. Asking over MCP recorded the deficit behind a refusal and asking on the CLI did not, so the same question had different consequences depending on where it was asked. A sandbox and a standing budget both widen what runs without a person and neither appeared in ambit authority or ambit status. ambit goal warned about a conflict with the alternative a proposal would not choose. Deficits captured from a runtime attribute to the provider that failed, which the curriculum could never see, so they roll up to what the provider supplies. An unseeded graph made the MCP server move every tool's answer under a nested key, changing the typed surface's shape according to the state of a database; the notice now rides on the text half only.
The tracker plugin honoured only one of the two database environment variables and then fell back to the session's working directory, wrote without the busy timeout the engine sets for four concurrent writers, and omitted kind — so every combo it created was stored as a provider and was invisible to half the product. Seven dead files are gone: two plugin aliases that re-exported names the canonical files already re-export, two component shims of the same kind, and three unreferenced scripts, one of which wrote a fixture that no longer exists.
#What makes the loop widen, not just report
The previous entry gave an agent a way to notice a gap, name it and ask. Nothing in that made the environment more capable: an agent could be told no all week and end it exactly as able as it started. Nine changes to the part that decides whether anything comes of the asking.
A threshold now counts successful work as well as passing checks, so an environment where the job succeeds daily stops having to run synthetic self-tests to earn trust it has already demonstrated. A failed run still counts against nothing, because attributing a run's outcome to every capability it touched would demote whatever a bad afternoon went near.
ambit authority promote with no arguments now names the grants that have earned a threshold nobody set. The interruption a threshold would end is exactly what stops anyone noticing it could, so the graph says it rather than waiting to be asked.
Scope decides. Two rules, in order: a forbidden grant wins outright at any specificity, and among the rest the most specific covering scope governs. Under narrowest-wins alone a grant saying "autonomous on staging" could never beat a standing "confirm everywhere", which made the only trade anyone actually wants inexpressible. --scope on a promotion writes that grant and leaves the standing one untouched. ambit authority sandbox declares somewhere acting does not matter, which is where a scoped threshold gets met cheaply; it relaxes confirmation and never a refusal.
ambit budget set <cap> --amount=$20 --by=<person> grants standing spend. Budgets existed and could only be written by the code that recorded spend, so the ceiling was real and the delegation was not.
ambit reject <id> <person> "why" records the half of every decision that used to vanish. ambit preferences --observed reads both halves, and ambit propose drafts the alternative the record favours and says why. A trait needs three decisions to count, and one that has gone both ways reads as contested rather than settled by majority.
ambit proposals --pending and the approval push carry the decision rather than announcing that one exists, and ambit approve takes several ids and one name. Every acquisition costs one interruption; a week of drafts read together is one sitting.
ambit reversible publishes what could be acquired without a person and what could not, which is the same list read backwards. ambit objects <target> and ambit verify --target= give authority and evidence an object, so committing to one repository forty times stops being a claim about the next one.
Five more MCP tools, twenty-two more tests.
#Ambit reaches the agent at the moment friction happens
Every command until now waited to be asked, which is no use to the agent that hits a missing binary mid-task, works around it, and hits the same one next week.
ambit briefing, served to a client on connect as the MCP resource ambit://briefing, says what is reached and proven, what is configured but failing, what is waiting on a person, what blocked work in the last week, and what is worth reaching next — prose, capped near 1,200 tokens, trimmed from the bottom because the order is the order of usefulness.
The ledger fills itself. The telemetry bridge reports failed tool executions and the engine classifies them from the signals a runtime states outright — a shell's message for a missing binary, EACCES, ECONNREFUSED, a 401, an MCP error kind — into the classes the deficit reports already used. A failure that only means work went wrong is left alone; one that cannot be attributed to any capability is kept and named, because that is a gap in the model rather than in the environment.
ambit_can now answers yes, ask or no with a sentence a person can act on, and files the deficit itself on a refusal, so asking before an unfamiliar tool costs one round trip rather than two. The README and llms.txt carry the two lines that make it a habit.
ambit next answers what to reach next and why — by what has actually blocked work once the ledger has observations, by leverage per hour of setup before then, and it says which. ambit record skill:<name> --provides= --verify= puts a skill the agent wrote on the map, refusing any registration without a read-only check and running it on the way in. ambit authority promote <cap> <action> --after=N --by=<person> is the person saying stop asking me once this has proved itself: the grant widens when the evidence arrives and narrows on a single failing check, with nobody asked. ambit sync export|import moves the graph and the ledger between machines, carrying no commands and no grants, so a rebuilt container gets its history back without importing someone else's permissions.
Seven MCP tools and one resource; twenty-seven new tests. The remaining piece is the npm publish, which is why every path in still starts with a checkout.
#The rest
- Every query names its row.
src/engine/rows.tsdeclares what a row of each table looks like, matchingschema.sqlplus the columnsmigrate.tsadds, and the handle'sallandgettake that type at the call. Sixty-odd reads that used to cast toanynow say which columns they select, which is how the compiler found a report reading agoal_idcolumn the proposals table does not have, and an intervention with no capability being used as a map key. - The client store is
ambitStore.ts. The guide said so; the file was a 12-line re-export oftoolchainStore.ts. The names now match the guide, the demo fixtures live instore/demo.tsso the store reads as the product's state rather than its sample data, and the old name stays as a shim. App.tsxis a shell. The viewport, hotkeys, first-run guide, toast, and the AG-UI stream are hooks; the top bar, welcome screen, guide card, and toast are components; the URL rule is a pure function the demo-landing test imports instead of copying. Two hundred and sixty lines where there were seven hundred, and the stream and key listener mount once instead of re-binding on every render.- The API server lives at
src/server/api.ts, beside its config, scanners, and tests, rather than at the repository root. - The roadmap opens with a status table: eleven sections, what each has built, and what each still lacks, so a reader does not have to find the "unbuilt" paragraph eleven times.
- Getting to a first run is shorter. The README opens with a four-row "try it" block (hosted demo, Homebrew, Codespaces, MCP registration), Homebrew is documented as an install path for the first time even though the tap has shipped every release, and a
.devcontainerseeds a graph and starts the map so Codespaces is a one-click checkout.ambit webno longer insists on Bun, which nothing else in the repository needs. New readers get a FAQ, aSUPPORT.md, a comparison table against tool-RAG, workflow graphs, and package managers, and anllms.txton the demo site for agents deciding whether to recommend the tool. Aserver.jsonandmcpNameare ready for the MCP registry once the npm package exists, and the launch kit lists the directories to submit to. A first issue or pull request gets a reply that says what CI will run. - The CLI groups forty-two commands under five nouns —
graph,plan,check,govern,report. Grouping is presentation, not a rename: every flat verb still dispatches, soambit impact xandambit graph impact xare the same command and nothing anyone has scripted stops working.ambit helpcovers a first session andambit help --allis the full surface. - The MCP server advertises each tool once. Every tool was listed twice, as
ambit_*and as the legacytt_*, sotools/listreturned 96 entries for 48 tools and spent about 3,600 tokens of every agent's context on duplicates. Att_name written before the rename still dispatches; it is no longer advertised. Results now carry MCPstructuredContentalongside the text block, so an agent reads a field rather than parsing a string. - An approval is only spendable on what it authorised.
verifyApprovalchecked that an artifact was intact, unexpired and signed by the recorded approver, and nothing else — so any approved proposal was a bearer token for any gated action, and an approval to install a linter would authorise a production deploy.approvalCoverscloses it: an approval is valid only for a capability that appears in the steps of the proposal that was approved. - CI measures coverage with a floor under it, and runs on Node 24 as well as 22.
- Autonomous Control Plane & Safety Interception: Active runtime interceptor (
src/control_plane/proxy.ts,src/control_plane/cli.ts) that enforces capability DAG prerequisites, lifecycle checks (degraded/broken), and authority policies before tool execution. AMBIT_BLOCKED_UNAUTHORIZED& State Invariance: Rogue or out-of-order agent tool executions (such as unapproved production rollouts) are blocked at the control plane boundary with exit code2, guaranteeingpre_state == post_state.- Human-in-the-Loop HMAC Remediation: Blocked executions generate structured remediation proposals and cryptographic HMAC challenges (
ambit approve <proposal-id> <person>), minting signed artifacts verified prior to state mutation. - OpenTelemetry Trace Instrumentation: Spans and structured event records capture DAG evaluations, missing authorization nodes, HMAC challenges, and verification receipts.
- Verified Incident Trace & Demonstration Suite:
- Published Incident Report:
docs/incidents/INCIDENT_TRACE_001.md - Automated suite:
src/control_plane/proxy.test.ts— eight cases in the suite CI actually runs, replacing four pytest cases that no job had ever executed. - 90-Second Reproducible Demonstration:
scripts/demo_incident_trace.py(demo:incident) anddocs/incidents/demo_intervention_trace.cast.
- Published Incident Report:
#0.4.1 — 2026-08-30
- Adoption and distribution hardening: current install instructions, Node-only CLI bootstrap, npm release automation, and the
ambitHomebrew command alias. - The npm package is prepared for publication with a compiled engine (
dist-cli/, built onprepack) and no runtime dependencies. It is not advertised as an install path untilambit-cliis actually published. A bareambitfrom Homebrew or a checkout seeds an empty graph before reporting instead of telling you to run a second command.--jsonruns on a cold graph stay pure JSON. ambit mcpruns the MCP server from an installed copy —claude mcp add ambit -- ambit mcpregisters the verified Homebrew command without knowing where Homebrew placed the package.ambit sharewrites a self-contained HTML snapshot of the map for posting wherever the person chooses — the file is built from an allow-list (name, kind, domain, era, state, lifecycle, edges), so commands, URLs, paths, descriptions, and economics cannot enter it; people render as "a person", and--redactreplaces every non-curated name with its category. Nothing is uploaded; writing the file locally is the whole command.- The agent loop closes for real, and the README shows it. The embeddings capability's local alternative now carries a declarative
config_patchand an authority declaration (execute: confirm), which makes it the first acquisition that runs the whole span: an agent records a deficit and drafts a proposal over MCP, a person approves and applies it, the graph re-seeds, and Local Embeddings unlocks through composition.scripts/demo-agent-loop.tsruns that loop against a fixture graph and renders the actual output asdocs/assets/agent-loop-demo.gif— a failing loop fails the recording rather than rendering a fiction. The propose note and MCP tool description no longer claim applying is unimplemented; they say whatambit applyactually refuses and why. - Evidence reaches the surfaces people actually look at.
ambit statusgains an evidence block — proven / unproven / failing counts, when a check last ran, and which unproven capabilities have a declared check waiting — and the map badges each reached capability: ✓ for a passing check, ! for a failing one, nothing for configured-but-unproven, with the detail panel saying it in words ("Check passed 2h ago" / "Configured — never verified"). The lifecycle machinery existed; it was visible only to someone who already knew to askambit verify --history. ambit weboutside a checkout explains that the visualizer needs one (and where the hosted demo is) instead of failing with bun's script-not-found.- First-run discovery now combines OpenCode, Claude Code, Cursor (
~/.cursor/mcp.json), Windsurf (~/.codeium/windsurf/mcp_config.json), Gemini CLI (~/.gemini/settings.json), Claude Desktop (its platform config path), and Codex CLI (~/.codex/config.toml, read by a purpose-built reader for its[mcp_servers.*]tables rather than a TOML dependency). Shared MCP servers stay one capability with a contribution edge from each runtime. - The README no longer advertises the unpublished npm package; the release workflow will publish it automatically once
NPM_TOKENis configured. - Redundancy now accounts for what fails together. Providers presenting the same credential do not fail independently, so a capability with three of them was reported as robust — and, since having several providers is what kept it out of the single-point-of-failure report, excluded by the very fact that made it fragile. A
credentialsblock declares which providers share one,ambit impactcalls what survivesnominalrather thanredundant, andambit credentialsanswers what a revocation would end. Only a credential every provider presents counts; partial sharing leaves something standing and is not reported as fragile. No secret is read or stored — there is no field one could arrive in. - Commands whose answer is a list printed the note about the list and not the list.
ambit authorityshowed its per-row detail and omitted the four lists summarising it, on the surface the tool considers primary. - Capability, action, provider, resource, actor and runtime are now separate kinds of node, and
provides,contributes,requires,optional,authorizesandruns_onseparate kinds of edge. Redundancy analysis previously matched three English sentences to decide what supplied what, so an adapter phrasing one differently made a capability with two providers report as a single point of failure. - Ids are unchanged, so no ledger history is lost. Existing graphs migrate by adding columns and backfilling once;
ambit spof,ambit impact,ambit planandambit sincereturn identical output against a graph seeded by the previous version. - Capabilities can declare the concrete actions they confer. Ten do, and
ambit actions version-controlreports that reading the repository and committing may run unattended while pushing a branch and merging to the default branch may not — a distinction the coarse node could not make. - Authority is a table with a source, and both adapters now pass through what their runtime states about itself: Hermes's
approvals.modeandapprovals.cron_mode, Claude Code'spermissions.defaultMode. Where the curated model and the runtime disagree the narrower wins, andambit authoritynames which one narrowed it. Nothing is enforced; Ambit describes authority. - Each capability carries a lifecycle — unknown, detected, configured, verified, reliable, degraded, broken — derived from its providers and its recorded evidence. A capability whose check has started failing reads as broken and stays in the frontier, because reachable and working are different claims — and it now gates:
ambit planrefuses to route through it,ambit simulatereports it asblocked_by_degraded,ambit authorityandambit actionsstop listing it as reachable,ambit nearsays re-verify rather than add, andambit statsreports it in a failing count. The ledger records demonstrated reliability beside reach: each snapshot carries a verified count and an id→lifecycle map, soambit sincereports a capability that stopped working as diminished while the structural frontier stays flat. ambit sincedistinguishes an expanding vocabulary from an expanding frontier. Upgrading to this version would otherwise have reported a dozen capabilities gained on a machine where nothing happened.- The engine is now modules named after what they do — discovery, inference, assurance, planning, governance, ledger — rather than one 1,864-line file.
- The graph an installed copy keeps now lives in
~/.local/share/ambit/graph.dbrather than inside the install directory, wherebrew upgradedeleted it. A checkout still keeps its graph in the checkout, and an existing graph is never moved.ambit whereprints the path. ambit seedandambit whereare documented inambit --help; the Homebrew and npm installs previously had no documented way to build a graph.- An unseeded graph now says so instead of answering every question with "Nothing to report". Over MCP it returns an explicit notice, so an agent cannot mistake not set up for an environment with no capabilities.
- The MCP server introduces itself as
ambitat the package version, rather thantech-treeat1.0.0. - CI runs typecheck, tests, build, and a bootstrap against a machine with no agent config. The demo deploys from CI.
#0.4.0 — 2026-08-13
Capability Graph became Ambit. The old name described the data structure; the new one describes the subject — what you, your agents and your machines can jointly do. Mostly a repair release: the previous version documented a tool that did not run, so the claims it already made had to become true before anything could be built on them.