Data Built for the Gap in Your Model
Our engineers work alongside yours, from the task design through to the data itself. We'll dig into the problem with you and come back with a pilot sample built to your spec.
We Measure It Ourselves
Coding RL environments
Refactoring and unit-test tasks as RL environments. A real repo snapshot in an isolated container, returning a graded reward rather than a pass mark.
- Calibrated to a 20 to 80% pass band
- Graders hardened against gaming attempts
- Rubrics grounded in Fowler and Ousterhout
TypeScript, Python, Go and Rust. Local renames through inheritance restructuring.
Computer use SFT
Desktop sessions captured as canonical episodes. One episode projects into SFT, grounding, preference, repair and evaluator sets.
- Cross-app planning with source-of-truth conflicts
- Long-horizon editing with verified staged artifacts
- Repair tasks scored against a hidden error list
Negative examples must be rejected by validators before a package ships.
MCP tool-use environments
Executable agent tasks, each with a restorable initial state and a verifier that scores the resulting environment rather than the model's prose.
- pass@5 up from 33.3% to 50.0%
- avg@1 up from 18.7% to 25.3%
- Open 9B model, one GRPO + LoRA epoch, five runs
Measured on a fixed 30-task set, filesystem subset only. Database tasks were not in this run.
The Areas We Go Deep In
Agentic, Tool Use & Long-Horizon
Tasks that teach a model to plan, reach for the right tool, and recover when a run goes sideways.
Computer Use & GUI Agents
Real desktop work, captured end to end: read the screen, act, confirm the state changed, verify the file.
RL Environments & Verifiable Rewards
Working systems rather than static files, with reward logic wired to the behavior you want to reinforce.
Coding
Repository-scale work from engineers who ship production code: build, refactor, test, optimize.
Multimodal & Visual Reasoning
Aligned image, video, audio and text, weighted to where frontier vision models still break.
Expert STEM Reasoning
Original Olympiad and PhD-level problems, written to probe the ceiling rather than test recall.
Long-Context & Document-Grounded Reasoning
Tasks where the answer lives in the attachments, and a rubric checks what comes back.
Evaluation, Rubrics & Diagnosis
Evaluations that find where a model breaks, and rubrics that make quality measurable.
SFT, RLHF & Preference Data
Demonstration, preference and critique data from people qualified to judge the domain.
Data We Already Hold, Ready to License
Exam & problem banks
University and K-12 STEM across several languages, manually annotated and OCR-processed.
Books, theses & academic papers
Textbooks through doctoral research, scanned and native format.
Game source code
Complete game repositories across Unity, Cocos Creator and native stacks, screened to exclude anything already public on GitHub.
Encyclopedic & web
Cleaned web-scale text and structured reference content.
Image-text corpora
Paired image and text with structured metadata, not loose captions.
Conversational & Q&A
Dialogue and Q&A, including consultation transcripts and bilingual speech text.
News & media archives
Journalism spanning print, broadcast and online publishing.
Professional knowledge
Material from regulated professions where accuracy carries consequences.
Code solutions
Competitive programming solutions with the discussion threads behind them.
