Book a Demo

Agent trajectory data · for AI labs

Real coding trajectories.
Not staged tasks.

Trajector licenses complete, consented Claude Code sessions from real development work — prompts, tool calls, diffs, failures, fixes, and commits — collected from open-source developers who get paid for it.

Book a Demo
session_7f3a.jsonl · claude-sonnet-4-6
00:00 user fix the flaky retry logic in the sync worker
00:04 tool_use read src/sync/worker.ts
00:19 tool_use edit worker.ts +14 −6
00:31 tool_use bash npm test
00:47 tool_result 2 failed — backoff never resets
01:02 tool_use edit worker.ts +3 −1
01:15 tool_use bash npm test
01:28 tool_result all 41 passed
01:44 commit a41f2c9 fix: reset backoff on successful sync

What’s inside every trajectory

Not snippets. The complete record of how a real problem got solved, as structured JSONL with full event ordering and model metadata.

user / prompts

The developer’s actual intent, in their own words, at every turn.

tool_use / tool_result

File reads, edits, and shell commands with their real outputs — including failures.

diffs & tests

Every change and the test results that accepted or rejected it.

commits & outcomes

Sessions linked to their commits, so trajectories carry ground-truth resolution.

Why natural beats manufactured

Most coding data on the market is produced for the transaction: preset tasks, performed for bounties, in artificial environments. Trajector captures work that would have happened anyway.

Staged task data

  • Tasks designed by the vendor, not the real world
  • Performed for the payout, under observation
  • Clean, linear solutions — few genuine dead ends
  • Distribution limited by task authors’ imagination

Trajector trajectories

  • Real problems from live codebases
  • Natural behavior — the incentive follows the work
  • Authentic failure, retry, and correction patterns
  • Distribution as wide as open source itself

Verified before it ever reaches you

Every session passes a multi-stage acceptance pipeline. What fails, you never see — and never pay for.

CAPTUREDevice-bound

Uploads signed with per-device keys from attested, reproducible CLI builds.

SCRUBSecrets removed locally

Keys, tokens, and env vars stripped on the contributor’s machine, verified server-side.

VALIDATEStructure & timing

Schema, event ordering, and temporal consistency checks against real API latency.

DEDUPEUnique sessions only

Exact and near-duplicate detection across the full corpus.

REVIEWCoherence-scored

Sampled sessions judged for genuine task resolution before corpus inclusion.

Rights you can build on

Trajectory data is only as usable as its provenance. Every session arrives with the consent chain intact.

Explicit contributor consent

Collection is opt-in per project. Nothing is captured from projects a contributor hasn’t enabled.

Rights attestation

Contributors attest that shared sessions cover code they own or that is open source, recorded at onboarding.

Commercial training license

Datasets are delivered with clear usage authorization for model training and evaluation.

Deletion honored end-to-end

Contributors can pause, exclude repos, or delete local captured data at any time.

Power your models with real trajectories

Start with a scoped pilot batch.

Book a Demo