Want help running the benchmark? Sign up to benchmark Teamwork Graph with your own work.
This benchmark answers one question: does connecting your AI agent to Teamwork Graph produce better answers, using fewer tokens, on your own work?
You run a fixed set of real prompts twice — once without Teamwork Graph, once with it — then pick the better answer in a blind review. The tool produces a report comparing the results side by side.
The recommended workflow runs through the Teamwork Graph AI Context Benchmark skill in Codex or Claude Code and takes about 20–40 minutes end to end. Your data never leaves your machine unless you explicitly choose to share a local report bundle.
This tool is in active development. Results are directional, not final. Read how to interpret your results before quoting any number — token savings and answer quality do not always move together.
To run the benchmark, you need:
| Stage | What happens |
|---|---|
| Prepare | Install the Benchmark CLI, sign in to your agent and Teamwork Graph CLI, and connect at least one code source if needed (GitHub or Bitbucket). |
| Validate | A readiness check confirms connectors and source reads — no tokens spent. |
| Run | The same prompts are answered twice: baseline vs Teamwork Graph. |
| Review | A browser opens; you pick the better answer for each prompt. |
| Report | Quality, token, and latency results, broken down by prompt. |
| Share | Optional: package and send results to Atlassian for analysis. |
The benchmark keeps the agent and prompts fixed so the Teamwork Graph connection is the only meaningful difference.
| Baseline arm | Teamwork Graph arm |
|---|---|
| Your agent uses standard Atlassian access, with Teamwork Graph turned off. | The same agent answers the same prompts with the Teamwork Graph CLI connected. |
The Teamwork Graph AI Context Benchmark skill drives the supported workflow: readiness, isolated baseline and Teamwork Graph lanes, blind review, and the report handoff.
Rate this page: