Labs
Everything we bring you, we ran on ourselves first.
Labs is where we test every new model on our own evals, and build our own products, automations and workflows. We add to it all the time, and only what earns its place reaches clients.
Evals
Every release, tested.
We build our own benchmarks and run every new model through them, so we know what each one is good at before we put it in front of a client.
AI evals for your company- Our own benchmarks, built on real work, not puzzles
- Every new model run through them the moment it ships (and often before)
- Where each model is strong, where it breaks, and what changed since the last release
- Results published here as we go
What we build
Always something new. We add to Labs every week.
Experiments
Every new model and tool, tried on real work the moment it ships (and often before).
Evals
Our own benchmarks, run on every new release.
Automations and workflows
Agents and workflows that run our own business, day to day.
Products
Software we build for ourselves first, and for clients once it's proven.
The tools we test
The models and tools we test the moment they ship (and often before), and run on our own business.
- OpenAI
- Anthropic
- Claude Code
- Codex
- Cursor
- GitHub
- Cognition
- Perplexity
- Grokbot
- SpaceXAI
- Hugging Face
- TypeSafe
Let us help.
Tell us where the hours go and what you've tried so far. We'll reply within one working day.