Every Coding Agent Trend Is Program Synthesis
Prompt engineering, context engineering, loop engineering, harness design, evals. If you know the source material you can predict which one gets named next.
On this page
Every coding agent trend is a shadow of a program synthesis primitive: named by people who can see the shape on the wall but not the object casting it. Watch the vocabulary arrive in order. “Prompt engineering” was circulating by July 2020 and pinned down in print by February 2021. “Evals” went from a repo OpenAI open-sourced the day GPT-4 shipped to a job title. “Context engineering” took over in June 2025, six days from Tobi Lütke’s tweet to Karpathy’s +1. This February, “harness engineering” was named and adopted inside a single week, and there are now separate eight-minute explainer videos for “loop engineering” and “agent harness”. And the topology itself is being named now. Agent work gets drawn as an explicit graph, nodes doing the work, edges carrying the results. LangGraph turned that framing into a framework in 2023-24, and Anthropic’s agent-patterns post cataloged the topologies: routing, parallelization, orchestrator-workers. Each term arrives as a new discipline, and each names a piece of machinery that program synthesis built decades ago.
I keep noticing this because I’ve spent 2026 running the classical loop at catalog scale with coding agents. The trends feel intuitive for a boring reason. They are one old loop, renamed a primitive at a time.
The loop, before it had product names
In 1969, Cordell Green’s QA3 system extracted working programs from resolution proofs. State what you want as a theorem, and the proof of it contains the program (Application of Theorem Proving to Problem Solving, IJCAI 1969). Manna and Waldinger built that into deductive synthesis, where the program falls out of a constructive proof of the specification (A Deductive Approach to Program Synthesis, TOPLAS 1980).
The version that matters most for agents showed up in 2006. Solar-Lezama and colleagues built a solver around “a counterexample-driven iteration over a synthesize-verify loop built from two communicating SAT solvers” (Combinatorial Sketching for Finite Programs, ASPLOS 2006). His 2008 thesis named the pattern: “a technique we have named counterexample guided inductive synthesis, or CEGIS” (Program Synthesis by Sketching, UC Berkeley 2008). Propose a candidate, verify it, and when verification fails, feed the counterexample back to sharpen the next proposal. If that sounds like every agent loop you’ve run this year, that’s the point.
Sumit Gulwani compressed the whole field into three dimensions in 2010: “expression of user intent, space of programs over which to search, and the search technique” (Dimensions in Program Synthesis, PPDP 2010). A year later he shipped the proof that this line of work ships products. FlashFill synthesizes string programs from input-output examples, and it’s in every copy of Excel (Automating String Processing in Spreadsheets Using Input-Output Examples, POPL 2011). And in 2013, SyGuS formalized the move everyone now performs daily: “allowing the user to supplement the logical specification with a syntactic template that constrains the space of allowed implementations” (Syntax-Guided Synthesis, FMCAD 2013). The 2017 survey by Gulwani, Polozov, and Singh is the single best on-ramp to all of it.
Intent, space, search, verify, feed the failure back.
The mapping
| The trend term | The synthesis primitive | The evidence |
|---|---|---|
| Prompt engineering | Expressing user intent (the specification) | Reynolds & McDonell 2021: prompting “may be conceived as programming in natural language” |
| Context engineering | Constraining the search space | Karpathy’s definition, “filling the context window with just the right information,” is SyGuS’s “syntactic template that constrains the space of allowed implementations,” twelve years on |
| Loop engineering / harness design | The search technique | Willison: “An LLM agent runs tools in a loop to achieve a goal” |
| Evals | The verifier | OpenAI, Mar 2023: “a framework for evaluating OpenAI models and an open-source registry of benchmarks” |
| Hooks / guardrails | Protecting the verifier from the candidate | DeepSeek-R1 skips a learned reward model because it “may suffer from reward hacking,” the problem named in 2016 |
| Autoresearch | The full CEGIS loop aimed at experiments instead of functions | Karpathy’s own recipe: “the human iterates on the prompt (.md),” “the AI agent iterates on the training code (.py)“ |
| Spec-driven development | The specification, named a second time | GitHub Spec Kit: “you start with a (you guessed it) spec”; Kiro ships the IDE around it |
| Verifiable rewards | The verifier, named on the research track | Tülu 3: “a novel method we call Reinforcement Learning with Verifiable Rewards” |
| Multi-agent orchestration | The search technique, parallelized | Anthropic: “a lead agent coordinates the process while delegating to specialized subagents that operate in parallel” |
| Agent memory | The state the loop carries between iterations | MemGPT: “virtual context management,” borrowed from OS memory hierarchies |
| Agent graphs | The search technique’s topology: candidate generation as a DAG, verifier gates on the edges | LangGraph plus Anthropic’s patterns: routing, parallelization, orchestrator-workers; named “Graph Engineering” in August 2026 |
| Decision models | The verifier, as a model: yes or no and how sure, but not why | The New Stack: they “return a set of predefined answers with confidence scores”; Jev lists “Verify everything” among its uses |
Specification: prompt engineering, then spec-driven development
Both name the same step: say what you want precisely enough that a candidate can be checked against it.
Search space: context engineering
Context engineering is search-space design. The SyGuS insight was that a logical spec alone leaves the search hopeless; a grammar over allowed implementations makes it tractable. A CLAUDE.md is a grammar over behaviors. SutroYaro’s search_space.yaml is one literally: it enumerates what an agent may mutate, and the harness rejects everything outside it. When I wrote in coding-agents-are-the-base-agent that most agent errors are map problems, this is the theory underneath. A bad map is an unconstrained search space, and unconstrained search was the failure mode synthesis spent the 2000s escaping. By mid-2025 the failure even had its own name. Context rot, Chroma’s report on model performance growing unreliable as input length grows, is the search-space problem measured from the inside.
Search technique: loops, harnesses, multi-agent orchestration, agent graphs
The agent itself is the search technique. Chimera’s decomposition (provider, tools, loop, environment) is a menu of search strategies over the same space. Racing N different loops on one task, which is what Chimera’s multiplexer does, is a search-technique bake-off. In synthesis terms, none of it is exotic. Enumerative vs. stochastic vs. deductive search is a settled taxonomy in the 2017 survey, and which loop wins on which problem is an empirical question you can benchmark, not a matter of taste.
Verifier: evals, verifiable rewards and decision models
Evals are the verifier, and the verifier is the part that decides whether any of it was real. Synthesis learned early that the verifier defines the product. Your program is only as correct as the check it passed. My most expensive 2026 lesson was the same lesson. Bourbaki’s v0.2.1 claimed 94.3% on miniF2F because the Lean REPL reported success on tactics that failed standalone compilation; the audited number was 6.2%, and the release is retracted with the inflation preserved for the record. CEGIS treats verification failure as signal to feed back. Agent practice mostly treats it as a grade. The feedback edge is the part of the loop the industry hasn’t renamed yet, which is why I run audits per wave instead of evals at the end.
Decision models are the checking step, sold as a model. In the 2006 loop the checker was a solver. It said whether the candidate was right, and when it wasn’t, it handed back an input that broke it. Logic has a name for a checker that always finishes with a correct yes or no, a decision procedure, so the name is older than the models. This September a decision category of models showed up that does only the yes-or-no part. TypeSafe released Jev on September 15 and calls it a System One model, and The New Stack calls the category decision models. They don’t write text. You give them the state and a question with the answers you allow, and they pick one and say how sure they are. Sebastian Raschka called Jev “the ChatGPT moment for classification.” TypeSafe lists “Verify everything” among the uses, and its FAQ says it needed a new way to train, because rewarding a model with a check written in code only works when a simple check exists, and most real-world judgment calls don’t have one. Two weeks later OpenAI announced its own Decisions API, pitched at sorting content, routing requests and picking an agent’s next step. Neither hands back the input that broke anything. You get an answer and a confidence number.
The first studies already put numbers on it. One looked at the 2,170 public Jev projects on GitHub a week after launch, labeled by LLM agents, and found 77% use it to judge something about the input, 52% to score or rank, and 31% to pick the next action. Another tested Jev as a judge. It comes within three points of GPT-6 when the answer can be read straight off the text, at 0.36% of GPT-6’s price, and falls behind when the answer has to be worked out, as in math, code and logic. A third found that Jev goes by an answer’s name more than by what the answer is defined to mean. Call the two options yes and no, swap the definitions behind them, and it flips 24 times as many answers as asking the same thing twice does, while every answer still comes back in the right format.
Protecting the verifier: hooks and guardrails
Hooks and guardrails keep the candidate away from the check. An agent with write access can edit the test instead of the code, and a hook that blocks writes to the tests is oracle protection at the size of one repo.
State between iterations: agent memory
Agent memory is the state the loop carries between iterations. In CEGIS that state is the pile of counterexamples, and each failed check narrows the next proposal. MemGPT arrived in 2023, borrowing virtual memory from operating systems.
The whole loop: autoresearch
Autoresearch is all of it at once, pointed at experiments instead of functions. In Karpathy’s recipe the human edits the prompt file and the agent edits the training code, so the spec and the candidate sit in separate files, and the training run is the check.
Why the names keep coming
The charitable read is the correct one. Builders hit each wall in order and name it on impact. The spec came first (prompt engineering, 2020-21). Checking came next (evals, March 2023). The material around the spec followed (context engineering, June 2025), then the iteration structure itself (agentic loops and harnesses, late 2025 into 2026), and last the assembled loop pointed at research (autoresearch, March 2026). Read against Gulwani’s dimensions, the arrival order walks the primitives: user intent, verifier, search space, search technique, and finally the whole CEGIS loop.
The renaming has also started its second lap. Specification got named twice, first as prompt engineering in 2020, then as spec-driven development in 2025, when Kiro shipped an IDE around it and GitHub shipped a toolkit. The verifier got named on two tracks at once. Evals covered the product side, with eval-driven development as the discipline form. On the research side, Tülu 3 coined Reinforcement Learning with Verifiable Rewards in November 2024, and DeepSeek-R1 made rule-based verifiable rewards the headline recipe two months later. DeepSeek’s stated reason for skipping a learned reward model, that it “may suffer from reward hacking,” points at the one primitive named before any of this. Reward hacking has been on the books since 2016. Memory ran the pattern in reverse. MemGPT arrived in 2023, no “memory engineering” coinage ever landed, and the practice got absorbed into context engineering instead. The search technique is the station renamed fastest, from the graph framing (LangGraph, 2023) to multi-agent orchestration (June 2025) to agentic loops (September 2025) to harness engineering (this February) to graph engineering (this August). One station, five names in three years, and the newest drawings put verifier nodes directly on the graph’s edges. Cohorts rename the stations as they reach them; the stations don’t move.
You can catch the naming happen live. Mitchell Hashimoto, this February: “I don’t know if there is a broad industry-accepted term for this yet, but I’ve grown to calling this ‘harness engineering.’” Six days later, OpenAI published a post titled “Harness engineering.” That’s the whole mechanism inside one week. Names help adoption, and adoption is good.
It cuts the other way too. The trend terms are so young they lack provenance. “Prompt engineering” gets retro-attributed to Gwern, who actually wrote “prompt programming”, and its earliest well-documented definitional use in print is Reynolds & McDonell, February 2021. The synthesis primitives have DOIs.
The practical difference for anyone who reads the old material: you stop waiting for the name. The primitives that haven’t trended yet are sitting in the same papers, visible now.
Three predictions, on record
If this lens is right, it should predict. Dated 2026-09-12, checkable later:
- Software factories are next. They are the epitome of spec to implementation.
- The critics of software factories say they lack verification. So the next set: agents that are good at verification.
- Eventually we move to new categories of super apps rather than super agents.
If any of these arrives under some fresh name, this table gets a new row and the thesis gets a data point. If none arrives, the lens is weaker than I think and that goes in the log too. That’s the deal I’ve signed up for everywhere else. Claims you can check beat claims you can admire.
Update, 2026-09-30. The next set showed up three days after I dated these, as a decision category of models rather than agents. Within 24 hours of Jev landing on Vercel’s AI Gateway, nearly 13% of its paid teams were using it. But OpenAI pitched its version at sorting and routing rather than checking, and the first study of Jev as a judge puts it behind on code, which is what software factories need checked. So prediction 2 is half there, and the table gets a new row.
The seat I’m watching from
The catalogs ran as a synthesis pipeline with the roles made explicit: spec as a GitHub issue, one candidate generator per stub (53 for Hinton, 58 for Schmidhuber), a read-only verifier per wave, human acceptance at the gate. SutroYaro points the loop at experiments. Novalis points it at a terminal. The vocabulary around all of this will keep changing. The loop hasn’t changed since Green extracted a program from a proof in 1969, and I expect it to outlast the next six names too.
Links
- Green 1969, Application of Theorem Proving to Problem Solving (IJCAI)
- Manna & Waldinger 1980, A Deductive Approach to Program Synthesis (TOPLAS)
- Solar-Lezama et al. 2006, Combinatorial Sketching for Finite Programs (ASPLOS)
- Solar-Lezama 2008, Program Synthesis by Sketching (UC Berkeley PhD thesis; coins CEGIS)
- Gulwani 2010, Dimensions in Program Synthesis (PPDP)
- Gulwani 2011, Automating String Processing in Spreadsheets Using Input-Output Examples (POPL; FlashFill)
- Alur et al. 2013, Syntax-Guided Synthesis (FMCAD)
- Gulwani, Polozov & Singh 2017, Program Synthesis (Foundations and Trends in Programming Languages)
- Li, Parsert & Polgreen 2024, Guiding Enumerative Program Synthesis with Large Language Models (CAV; an LLM inside a SyGuS enumerator)
- Reward hacking: named in Concrete Problems in AI Safety, 2016
- Prompt engineering: circulating on Hacker News, Jul 2020 · first definitional use in print, Feb 2021 · LMQL, “Prompting Is Programming”, PLDI 2023
- Evals: OpenAI Evals, open-sourced the day GPT-4 shipped, Mar 14 2023
- Agent memory: MemGPT, Oct 2023
- Agent graphs: LangGraph, 2023 · Anthropic’s topology catalog, Dec 2024
- Graph engineering: named in arXiv 2608.21156, Aug 21 2026, whose abstract lists prompt, context, harness and loop engineering as its predecessors
- Verifiable rewards: coined in Tülu 3, Nov 2024 · DeepSeek-R1’s rule-based recipe, Jan 2025
- Eval-driven development: evaluation as a continuous governing function, Nov 2024
- Context engineering: Lütke’s tweet, Jun 19 2025 · Karpathy’s +1, Jun 25 2025 · Anthropic’s definition, Sep 2025
- Multi-agent orchestration: Anthropic’s orchestrator-worker system, Jun 2025
- Spec-driven development: Kiro, Jul 2025 · GitHub Spec Kit, Sep 2025
- Context rot: Chroma’s report, Jul 2025
- Agentic loops: Willison on designing them, Sep 2025
- Harness engineering: coined Feb 5 2026 · adopted by OpenAI six days later
- Counterexample feedback with an LLM learner: Liu, Sala, Reps & Murali, Jun 2026 · TRIM, minimizing the residue of agent search, Jul 2026
- Autoresearch: the repo, Mar 6 2026 · the announcement, Mar 7 2026
- Software factories: Greenfield & Short’s book of that name, Wiley, Sep 2004 · StrongDM, “Software Factories and the Agentic Moment”, Feb 6 2026 · Ostrovsky, “Software Factories in September 2026”, Sep 2 2026
- The verification critique: Stanford CodeX, “Built by Agents, Tested by Agents, Trusted by Whom?”, Feb 8 2026 · METR, “Many SWE-bench-Passing PRs Would Not Be Merged into Main”, Mar 10 2026 · The Verification Horizon, Jun 24 2026: “generating complex candidate solutions is no longer difficult — reliably verifying them has become the harder problem”
- Verifier agents, state of play: Qodo’s $70M round as “the independent verification layer”, Mar 30 2026 · Qodo’s agent-to-agent code review, Sep 9 2026 · VeriCodeGen, NeurIPS 2026 workshop · a contract-driven adversarial verification harness, May 25 2026
- Decision models: TypeSafe’s Jev, Sep 15 2026 · Vercel on Jev’s first day on AI Gateway, Sep 18 2026 · OpenAI’s Decisions API, Sep 29 2026 · The Decoder’s DevDay report, Sep 29 2026 · Raschka, from bag-of-words to Jev, Sep 29 2026
- Decision models, measured: Jev in the Wild, 2,170 public projects, Sep 24 2026 (labeled by LLM agents) · JEV-as-a-Judge, Sep 22 2026 · Type-Safe Is Not Error-Free, Sep 22 2026
- Decision models, before the name: Kroening & Strichman, Decision Procedures, first edition 2008 · OpenAI’s Moderation endpoint, scores on categories OpenAI picked, Aug 10 2022 · UniEval, grading text with yes/no questions, Oct 2022 · RLCR, training models to say how sure they are, Jul 2025
- Super agents vs super apps: Axios, “Ph.D.-level AI super-agent breakthrough expected very soon”, Jan 19 2025 · OpenAI’s “ChatGPT: H1 2025 Strategy” memo, the super-assistant plan · Webster, “Smart Agents Replace Super Apps”, Jan 8 2026 · TechCrunch, “OpenAI is still working on that ‘super app’”, Jun 7 2026