Harbor-native
Families are Harbor task directories — the Terminal-Bench 2 contract your harness already executes. No custom runner, no import step.
Teraport Atlas authors, hardens, and certifies agent tasks in the open Harbor format — the standard your evaluation stack already runs. Every family ships with a machine-checkable acceptance certificate.

Families are Harbor task directories — the Terminal-Bench 2 contract your harness already executes. No custom runner, no import step.
A family layer generates seeded variants of every task, in the spirit of the METR Task Standard — fresh instances on demand.
Static lint, oracle and no-op calibration, determinism replay — run on every version, with the verdict recorded per stage.
Every run emits a content-hashed acceptance certificate — including rejections. Acceptance criteria are agreed before authoring, and the certificate is the evidence they are judged against.
Authored, evaluated, and audited in one place: author families with live validation, run evaluation packs across models, and drill from an aggregate score to the exact rollout transcript that produced it.

Atlas packages evaluation work in formats your stack can inspect, reproduce, and run without a custom runner. The task family layer builds on existing standards for agent environments, benchmark execution, and rubric-based scoring.
The family layer's parameterized tasks follow the spirit of METR's Task Standard, the reference for packaging agentic tasks with explicit environments and scoring.
Atlas families are Harbor task directories — the exact contract Terminal-Bench 2 executes. If your harness runs Harbor tasks, it runs Atlas families unchanged.
Optional weighted rubrics score partial credit the way PaperBench's rubric trees grade paper reproduction — graded steps instead of all-or-nothing rewards.
Short answers to what evaluation teams usually ask before putting a family in front of their stack.
No. Families are Harbor task directories — the Terminal-Bench 2 contract. If your evaluation stack executes Harbor tasks, it executes Atlas families today, with no custom runner and no import step.
It asserts what each gauntlet stage returned, against which content hash, at which thresholds: static lint, oracle and no-op calibration, and determinism replay. It records the stages that ran and their verdicts — nothing more. It is emitted for rejections as well as acceptances, so the record is the same whether we passed or failed.
Acceptance criteria are agreed in the contract before authoring starts, and the certificate is the evidence they are judged against — portable, machine-checkable, and tied to the exact content hash you evaluated. If a family misses the agreed bar, the remedy is written into the engagement rather than left to argument after delivery.
Reference solutions and tests are bind-mounted only at verification time, the verifier log directory is reset before grading, and families regenerate fresh variants per seed — so agents cannot read their grader, forge rewards, or memorize a fixed task instance.
Tell us the capability you need to measure and we scope a calibrated family against it, delivered in your harness. Engagements start with a scoped pilot rather than a self-serve account; workspace access is provisioned per organisation as part of it.
Βοηθάμε ομάδες στην Ελλάδα να ενσωματώσουν AI, να δημιουργήσουν AI-powered προϊόντα, να αυτοματοποιήσουν κρίσιμες ροές εργασίας και να εκσυγχρονίσουν τα συστήματα που τα στηρίζουν.