Ask a model what it knows about a fast-moving tool and you get a confident average of the internet as of a training cutoff. So the fleet does something else: it curates fifty authority videos, hands them to NotebookLM, and keeps only answers that arrive with citations.[03] The reading happens on Google's machines, not in the agent's context window. The distillate stays forever.
A language model's memory is a polite, plausible average. It is wrong in exactly the places that matter: the flag names, the limits, the current practice.
Every stage of software work now runs through agents, and agents answer from their weights unless you stop them. For settled knowledge that is fine. For a tool that shipped last month it is a confident guess dressed as fact.
Forge exists to remove that failure mode. You choose the sources. The model is only allowed to speak from those sources, and every claim it makes must carry a citation back to one of them.[03] What comes out is not a vibe. It is a cited distillate you can keep, audit, and build on.
Uncited means the model answered from its weights: the exact failure this pipeline exists to prevent.
The distill gate · forge/readme[03]
A sweep of YouTube across 25 to 40 query angles, ranked by authority rather than views. A human reads the candidate list and chooses the fifty. The list is the judgment.
NotebookLM reads all fifty on Google's machines and answers only from them, citing as it goes. Zero cost to the agent's context window. An answer with fewer than two citations is a hard fail.
The distillate is written into skills the fleet loads on demand. Each skill is then graded against a quiz the corpus itself wrote. Nobody marks their own homework.
Sweeps YouTube across dozens of query angles and ranks candidates by authority, not views. Output is a candidate list that a human reads and picks from by hand.[03]
Idempotent, resumable load of the fifty into a fresh notebook. Re-run it and it does only the missing work.
Hard gateExits non-zero unless every source is present, ready, and above a word floor. A video with no transcript is caught as thin, not counted as a win.
Runs the question set against the notebook, each question in its own conversation so answers never contaminate each other.
Hard gateAn answer with fewer than two citations is a fail, not a warning. Uncited means the model answered from its weights.
The distilled files are grouped into a few skills with deep references. They carry frontmatter fields, hook event names, and specific numbers that no model produces from memory.[03]
NotebookLM generates a quiz from the corpus, and the critic grades the skill against it.
Hard gateThe eval set is written by the corpus, so the skill author can never mark their own homework.
Two corpora running concurrently corrupt each other's answers. Notebooks run serially, and every question is wrapped in a timeout.[02]
The best SwiftUI animation channel alive ingested at 5 to 38 words of transcript per video. Beautiful to watch, worthless to a corpus. The word floor caught all four.[02]
They are rejected deterministically the moment they are added. Replace them with a narrated source. Do not retry.[02]
A keepalive job refreshes the session every two hours, then proves it works with a live check. Verifying the agent is loaded is not verifying the agent works.[02]
Forged claude-skill-forge · claude-orchestration · claude-context-discipline
Forged interface-craft · motion-craft · web-that-converts
Forged apple-design-language · swiftui-motion · swiftui-components
Forged sexy-ui · quality-bar · expensive-surface[01]