Skip to content
Products

Data Built for the Gap in Your Model

Our engineers work alongside yours, from the task design through to the data itself. We'll dig into the problem with you and come back with a pilot sample built to your spec.

Recent Work

We Measure It Ourselves

Coding RL environments

Refactoring and unit-test tasks as RL environments. A real repo snapshot in an isolated container, returning a graded reward rather than a pass mark.

  • Calibrated to a 20 to 80% pass band
  • Graders hardened against gaming attempts
  • Rubrics grounded in Fowler and Ousterhout

TypeScript, Python, Go and Rust. Local renames through inheritance restructuring.

Computer use SFT

Desktop sessions captured as canonical episodes. One episode projects into SFT, grounding, preference, repair and evaluator sets.

  • Cross-app planning with source-of-truth conflicts
  • Long-horizon editing with verified staged artifacts
  • Repair tasks scored against a hidden error list

Negative examples must be rejected by validators before a package ships.

MCP tool-use environments

Executable agent tasks, each with a restorable initial state and a verifier that scores the resulting environment rather than the model's prose.

  • pass@5 up from 33.3% to 50.0%
  • avg@1 up from 18.7% to 25.3%
  • Open 9B model, one GRPO + LoRA epoch, five runs

Measured on a fixed 30-task set, filesystem subset only. Database tasks were not in this run.

What We Build

The Areas We Go Deep In

Agentic, Tool Use & Long-Horizon

Tasks that teach a model to plan, reach for the right tool, and recover when a run goes sideways.

Computer Use & GUI Agents

Real desktop work, captured end to end: read the screen, act, confirm the state changed, verify the file.

RL Environments & Verifiable Rewards

Working systems rather than static files, with reward logic wired to the behavior you want to reinforce.

Coding

Repository-scale work from engineers who ship production code: build, refactor, test, optimize.

Multimodal & Visual Reasoning

Aligned image, video, audio and text, weighted to where frontier vision models still break.

Expert STEM Reasoning

Original Olympiad and PhD-level problems, written to probe the ceiling rather than test recall.

Long-Context & Document-Grounded Reasoning

Tasks where the answer lives in the attachments, and a rubric checks what comes back.

Evaluation, Rubrics & Diagnosis

Evaluations that find where a model breaks, and rubrics that make quality measurable.

SFT, RLHF & Preference Data

Demonstration, preference and critique data from people qualified to judge the domain.

Off-the-Shelf Data

Data We Already Hold, Ready to License

Exam & problem banks

University and K-12 STEM across several languages, manually annotated and OCR-processed.

Books, theses & academic papers

Textbooks through doctoral research, scanned and native format.

Game source code

Complete game repositories across Unity, Cocos Creator and native stacks, screened to exclude anything already public on GitHub.

Encyclopedic & web

Cleaned web-scale text and structured reference content.

Image-text corpora

Paired image and text with structured metadata, not loose captions.

Conversational & Q&A

Dialogue and Q&A, including consultation transcripts and bilingual speech text.

News & media archives

Journalism spanning print, broadcast and online publishing.

Professional knowledge

Material from regulated professions where accuracy carries consequences.

Code solutions

Competitive programming solutions with the discussion threads behind them.