Agent trajectory data · for AI labs
Real coding trajectories.
Not staged tasks.
Trajector licenses complete, consented Claude Code sessions from real development work — prompts, tool calls, diffs, failures, fixes, and commits — collected from open-source developers who get paid for it.
Book a DemoWhat’s inside every trajectory
Not snippets. The complete record of how a real problem got solved, as structured JSONL with full event ordering and model metadata.
The developer’s actual intent, in their own words, at every turn.
File reads, edits, and shell commands with their real outputs — including failures.
Every change and the test results that accepted or rejected it.
Sessions linked to their commits, so trajectories carry ground-truth resolution.
Why natural beats manufactured
Most coding data on the market is produced for the transaction: preset tasks, performed for bounties, in artificial environments. Trajector captures work that would have happened anyway.
Staged task data
- Tasks designed by the vendor, not the real world
- Performed for the payout, under observation
- Clean, linear solutions — few genuine dead ends
- Distribution limited by task authors’ imagination
Trajector trajectories
- Real problems from live codebases
- Natural behavior — the incentive follows the work
- Authentic failure, retry, and correction patterns
- Distribution as wide as open source itself
Verified before it ever reaches you
Every session passes a multi-stage acceptance pipeline. What fails, you never see — and never pay for.
Uploads signed with per-device keys from attested, reproducible CLI builds.
Keys, tokens, and env vars stripped on the contributor’s machine, verified server-side.
Schema, event ordering, and temporal consistency checks against real API latency.
Exact and near-duplicate detection across the full corpus.
Sampled sessions judged for genuine task resolution before corpus inclusion.
Rights you can build on
Trajectory data is only as usable as its provenance. Every session arrives with the consent chain intact.
Collection is opt-in per project. Nothing is captured from projects a contributor hasn’t enabled.
Contributors attest that shared sessions cover code they own or that is open source, recorded at onboarding.
Datasets are delivered with clear usage authorization for model training and evaluation.
Contributors can pause, exclude repos, or delete local captured data at any time.