Tag: Fieldframe
All the articles with the tag "Fieldframe".
-
The Hybrid Agent: Compiling a Human Expert into an ARC-AGI-3 Solver, and the Machinery That Keeps It Honest
How the Fieldframe program moved from knowledge benchmarks to ARC-AGI-3: an LLM mind driving a deterministic Python body, with a specific human expert's solving cognition compiled into its knowledge layer, and a verification system that makes every claim the agent produces independently checkable.
-
The Failure Mode Taxonomy: 13 Ways Frontier Models Reason Badly, and How to Catch Each One
Thirteen mechanism-level ways frontier models reason badly, split into score-affecting and metadata-only, each with a way to spot it and a countermeasure class. A field guide from a year of cross-architecture evaluation.
-
The Research Behind the HLE Score: A Year of AI Behavioral Research
The methodology behind the agent, the failure modes it catches, the products that came out of the same research moat, and where the program goes next.
-
51.85% on Humanity's Last Exam: How a Solo Researcher Built a Multi-Agent HLE Submission
1,119 out of 2,158 on canonical HLE. Single workstation, no GPU cluster, no fine-tuning. The architecture, the numbers, and what's next.