The Tax Agent That Teaches Itself: How Codex Is Rewriting Professional Compliance
OpenAI, Thrive, and Crete have built a self-improving AI tax agent powered by Codex — and it's not just faster than human preparers, it gets smarter with every filing.
Written by OutOfToken AI
June 5, 2026 · 4 min read · Synthesized from reporting by OpenAI Blog · How this works
Tax compliance is one of the most document-dense, rule-bound, and error-sensitive domains in professional services — which makes it either the worst or best place to deploy an autonomous AI agent, depending on your tolerance for risk. OpenAI, working alongside partners Thrive and Crete, has apparently decided it's the latter. The three organizations have built a self-improving tax agent on top of Codex, OpenAI's code-generation model, that doesn't just automate filings — it iteratively refines its own reasoning through feedback loops designed to push accuracy upward over time.
The Problem With Tax at Scale
Professional tax preparation is a workflow nightmare. A single corporate filing can require synthesizing hundreds of documents — W-2s, 1099s, K-1s, depreciation schedules, foreign income disclosures — against a tax code that changes annually and varies by jurisdiction. Human preparers are expensive, bottlenecked during filing season, and prone to inconsistency across large document sets. Existing software automates data entry but stops well short of adaptive reasoning. The gap between 'filling out a form' and 'understanding the tax implications of a complex transaction' has historically required a licensed professional. That gap is exactly where this Codex-powered agent is designed to operate.
A Three-Part Loop: Generate, Evaluate, Improve
The architecture behind the agent centers on what the team describes as a three-part iterative loop. First, Codex generates structured outputs — code, structured data queries, compliance logic — from unstructured tax documents and natural-language instructions. Second, an evaluation layer scrutinizes those outputs against known regulatory requirements and historical filing patterns, flagging inconsistencies or low-confidence determinations. Third, the system uses that feedback signal to refine its future responses, gradually tightening its accuracy on the specific tax scenarios it encounters most frequently. This isn't fine-tuning in the traditional sense — it's closer to a closed-loop reinforcement mechanism baked into the production workflow itself, allowing the agent to specialize without requiring manual retraining cycles.
"The agent doesn't just automate what humans already do — it builds a compounding advantage: every filing it processes makes the next one more accurate, faster, and cheaper to verify."
What Thrive and Crete Actually Built
Thrive and Crete brought the domain expertise and real-world filing volume that transforms a compelling demo into a production system. The partnership allowed the team to train the agent on actual compliance workflows, stress-test it against edge cases that generic benchmarks miss, and wire its outputs into existing professional review pipelines rather than replacing them wholesale. The result is a system that accelerates practitioner workflows rather than bypassing them — the agent handles document analysis, research synthesis, and preliminary form population, while licensed professionals retain oversight on final submissions. This human-in-the-loop design is both a regulatory necessity and a trust-building strategy, positioning the technology as augmentation before it earns autonomy.
The implications here extend well beyond tax season. A self-improving agent architecture that works in the high-stakes, regulation-dense world of tax compliance is effectively a proof-of-concept for every other professional domain that has resisted automation on the grounds of complexity — legal research, audit, clinical documentation. If Codex can teach itself to navigate the Internal Revenue Code with measurable accuracy gains over time, the question isn't whether AI agents will reshape knowledge work. It's how fast the feedback loops will close, and who controls them when they do.
Editorial Note
OpenAI's official blog is a reputable primary source for announcements about their products and partnerships. Codex was a real OpenAI model (though superseded by GPT-3.5/GPT-4) used for code generation tasks. However, the specific claims about tax automation accuracy and workflow improvements would require verification of the actual case study details and partner company results.
Claim Tracker
AI-assessed
No publicly documented evidence of this specific collaboration or product deployment found in major sources
Codex was indeed OpenAI's code generation model, though it was deprecated in 2023
This is an accurate characterization of complex corporate tax preparation
No specific technical details or evidence provided about how this self-improvement mechanism works
Modern tax software incorporates various AI/ML capabilities; this claim oversimplifies current market capabilities
Ask AI about this story
// discussion
sign in to join the discussion