See our results

Gemma 4 12BOriginal LLM vs. CB LLM

Same model, same GPU, and the same four serving workloads, measured side by side.

Efficiency Frontier

Comparing the original BF16 teacher to our CB LLM and a variety of LLM compression methods.

Measured scaling

Wins increase with model size.

The CB Converter has one engine that runs Gemma 4 2B, 4B, and 12B. As model size increases, the measured performance advantage of our CB technology grows with it.

Validated beyond one family.

CB conversion extends beyond Gemma. See the separately measured Qwen3 4B result below, followed by the growing set of models we have validated through conversion.

Measured internally. Independent reproduction is underway.

Completed
  • Single-model Gemma 4 results across 2B, 4B, and 12B.
  • One shared CB LLM runtime across all three Gemma representations.
  • Locked full five-shot MMLU and eight-process serving measurements.
Preparation underway
  • Building the disclosed Gemma 4 protocol and evidence package for independent execution.
  • Coordinating the next phase of third-party reproduction.
  • Underway: independent validation and an academic paper documenting the results.
Planned
  • Datacenter-class GPU testing and serving curves.
  • Customer-workload validation through pilots.
  • Expanded public reproducibility materials.

Customer-defined proof

Test one model against the quality bar that matters.

Run a scoped comparison on your workload before making an adoption decision.

Explore a Pilot