OsakaLondon

Labs

Everything we bring you, we ran on ourselves first.

Labs is where we test every new model on our own evals, and build our own products, automations and workflows. We add to it all the time, and only what earns its place reaches clients.

01 · Lab only2402 · Internal use603 · Early clients304 · Ready to scale1

Evals

Every release, tested.

We build our own benchmarks and run every new model through them, so we know what each one is good at before we put it in front of a client.

AI evals for your company
  • Our own benchmarks, built on real work, not puzzles
  • Every new model run through them the moment it ships (and often before)
  • Where each model is strong, where it breaks, and what changed since the last release
  • Results published here as we go

What we build

Always something new. We add to Labs every week.

  • Experiments

    Every new model and tool, tried on real work the moment it ships (and often before).

  • Evals

    Our own benchmarks, run on every new release.

  • Automations and workflows

    Agents and workflows that run our own business, day to day.

  • Products

    Software we build for ourselves first, and for clients once it's proven.

The tools we test

The models and tools we test the moment they ship (and often before), and run on our own business.

  • OpenAI
  • Anthropic
  • Claude Code
  • Codex
  • Cursor
  • GitHub
  • Cognition
  • Perplexity
  • Grokbot
  • SpaceXAI
  • Hugging Face
  • TypeSafe

Let us help.

Tell us where the hours go and what you've tried so far. We'll reply within one working day.

What do you need help with?
Team size
Annual revenue