Skip to content

You don't need a hyperscaler

Hyperscalers are expensive because they solve expensive problems. They train frontier models on giant GPU clusters and serve millions of users every day.

Using a hyperscaler for a 200-person company is like renting a temple when you need a workshop. Unless you sell models or AI at internet scale, you probably don’t need one.

Most teams don’t need to train a model. Serving inference to millions is a different business. Most businesses are not in the business of selling tokenized intelligence.

A team of 200 mainly needs AI for two things:

  1. Help the team build products and run the business.
  2. Make customers happier by building AI into the product.

Both depend more on inference than training. A company can serve this work on hardware it owns. Capacity is not infinite, but it does not need to be.

Unless you train models or research new transformer architectures, one well-built inference server can do the job.

Here is the back-of-the-envelope math.

Start with the workload

Take Qwen3.8-27B, a dense open-weight model with 27 billion parameters. It can handle coding, documents, images, tool use, and ordinary knowledge work. It also lets you disable thinking per request.

That last detail matters. Thinking mode is useful when a problem deserves a long chain of reasoning. It is wasteful for extraction, classification, rewriting, routing, support replies, and many tool calls. With thinking disabled, the model produces the answer without first generating a long reasoning trace. Users see lower latency, and the server spends its token budget on output they asked for.

Now put the model on eight NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. Each card has 96 GB of GDDR7 memory and up to 1.6 TB/s of memory bandwidth. The server has 768 GB of GPU memory and up to 12.8 TB/s of aggregate memory bandwidth.

That is a serious inference machine. It is still a machine your team can own.

Memory bandwidth sets the pace

Single-user token generation is often limited by how quickly the GPU can read model weights from memory. Qwen3.8-27B has 27 billion language-model parameters. At 4-bit quantization, those raw weights take roughly 14 GB before metadata and runtime overhead. The model fits comfortably on one 96 GB card, with room for its vision encoder, context, and cache.

This lets the server run one model replica per GPU instead of splitting every request across all eight cards. Replicas are simpler to operate and give the scheduler eight independent places to send work.

The useful number is aggregate throughput. With a good 4-bit runtime, continuous batching, sensible context limits, and thinking disabled, 600 or more generated tokens per second across eight cards is a reasonable deployment target. It is not a universal benchmark. Quantization format, prompt length, cache pressure, batch size, and serving software all move the result. Measure the exact stack before buying hardware.

The capacity is easier to understand in human terms. At 600 generated tokens per second, the server can produce 2.16 million tokens per hour. If an average response is 500 tokens, that is about 4,300 responses per hour. Two hundred employees could each make more than 20 such requests in that hour before generation throughput became the bottleneck.

People do not all prompt at once, all day. A scheduler absorbs bursts. Short jobs finish quickly. Long jobs queue. Hard reasoning tasks can still go to a frontier API when they earn the cost.

For a company of 200 people, this is enough capacity for a large share of daily AI work.

The economics have changed

An API looks cheap because it starts at zero. The bill grows with use, and the dependency grows faster than the bill.

The provider controls price, rate limits, model retirement, data policy, and acceptable use. A product built entirely on that API inherits every one of those decisions. Procurement can negotiate a contract. It cannot remove the dependency.

Owned inference has different costs. You buy the server, power it, cool it, monitor it, and keep someone responsible for it. Those are real obligations. They are also legible. The hardware is a fixed asset. Spare capacity can serve another model. The system keeps working when an API changes its terms.

This does not require an ideological ban on cloud models. Use them where they are clearly better. Send the hardest prompts to the strongest model available. Keep a fallback for demand spikes. Rent GPUs during a launch if the traffic calls for it.

But routine inference should not default to a remote API merely because that was the only practical option three years ago.

What you need instead

You need a model that performs well on your own evals. You need an inference runtime, a request queue, access controls, logs, monitoring, and a fallback route. You need someone who understands the system well enough to fix it.

That is infrastructure work, but it is ordinary company-scale infrastructure work. It does not require a datacenter campus, a power plant, or a billion-dollar agreement.

The distinction is simple. Frontier training still belongs to organizations with enormous clusters. Useful inference no longer does.

For a 200-person company, eight RTX PRO GPUs can serve a capable 27B model at hundreds of tokens per second while keeping the weights, prompts, and operating decisions under company control. Hyperscalers remain useful suppliers. They are no longer critical infrastructure for every AI product.

Own the common path. Rent the exceptional one.

© 2026 Marvin Danig. All rights reserved.