Harness-agnostic
Families are self-contained task directories with explicit environments and scoring, built to the contract your harness already executes. No custom runner, no import step.
Teraport Atlas authors, hardens, and certifies agent tasks in an open, harness-agnostic task format — one your evaluation stack can run as-is. Every family ships with a machine-checkable acceptance certificate.
Families are self-contained task directories with explicit environments and scoring, built to the contract your harness already executes. No custom runner, no import step.
A family layer generates seeded variants of every task, in the spirit of the METR Task Standard — fresh instances on demand.
Static lint, oracle and no-op calibration, determinism replay — run on every version, with the verdict recorded per stage.
Every run emits a content-hashed acceptance certificate — including rejections. Acceptance criteria are agreed before authoring, and the certificate is the evidence they are judged against.
Authored, evaluated, and audited in one place: author families with live validation, run evaluation packs across models, and drill from an aggregate score to the exact rollout transcript that produced it.



Atlas packages evaluation work in formats your stack can inspect, reproduce, and run without a custom runner. The task family layer builds on existing standards for agent environments, benchmark execution, and rubric-based scoring.
Short answers to what evaluation teams usually ask before putting a family in front of their stack.
No. Families are self-contained task directories with explicit environments and scoring. If your evaluation stack executes standard agentic task directories, it executes Atlas families today, with no custom runner and no import step.
It asserts what each gauntlet stage returned, against which content hash, at which thresholds: static lint, oracle and no-op calibration, and determinism replay. It records the stages that ran and their verdicts — nothing more. It is emitted for rejections as well as acceptances, so the record is the same whether we passed or failed.
Acceptance criteria are agreed in the contract before authoring starts, and the certificate is the evidence they are judged against — portable, machine-checkable, and tied to the exact content hash you evaluated. If a family misses the agreed bar, the remedy is written into the engagement rather than left to argument after delivery.
Reference solutions and tests are bind-mounted only at verification time, the verifier log directory is reset before grading, and families regenerate fresh variants per seed — so agents cannot read their grader, forge rewards, or memorize a fixed task instance.
Tell us the capability you need to measure and we scope a calibrated family against it, delivered in your harness. Engagements start with a scoped pilot rather than a self-serve account; workspace access is provisioned per organisation as part of it.
We help US businesses turn AI ambition into working software through product engineering, automation, system design, and delivery support.