AI

Together AI’s Platform Puts Generative AI Infrastructure and Pricing Under Scrutiny

Together AI’s Platform Puts Generative AI Infrastructure and Pricing Under Scrutiny

Pull up Together AI's pricing page and try to answer one plain question: what does it cost? You will find dozens of answers. There is a per-token price for each model, an hourly rate for each GPU, and a calculator that assumes your servers never sleep. That sprawl is why Together AI is under scrutiny now. In July 2026 the company raised $800 million at an $8.3 billion valuation, and it now sells everything from cheap open-model API calls to reserved H100 clusters. Buyers and investors want to know whether its low unit prices hold up once utilization, engineering labor and multi-year compute commitments are counted.

The short version

Together AI has grown from a model-hosting API into a full-stack infrastructure provider. Its usage-based pricing is published in detail, but the real cost depends on which of several services you pick and how steadily you use it. Its headline numbers are mostly company-reported, including bookings, developer counts, compute commitments and speed claims. Each one measures something different, so read them separately.

The momentum: what changed in summer 2026

The trigger was money. Together AI announced the Series C on its blog, with Aramco Ventures, NVIDIA, Vista Equity Partners and General Catalyst among the backers. TechCrunch put the valuation at $8.3 billion. That is up from roughly $3.3 billion at the Series B about 16 months earlier, a jump of around two and a half times.

Two other numbers came with the round. Investors separately committed more than 500 megawatts of compute, Together said. TechCrunch reported that the company logged annual bookings above $1.15 billion in its most recent quarter.

Neither figure means what a quick reader might assume. A megawatt commitment is a promise about future power and hardware. It is not capacity that is already serving tokens. Bookings are contracted future spend, not recognized revenue and not profit. I have watched the same conflation happen with CoreWeave's $100B backlog: the large forward-looking number gets read as current scale.

Together AI stopped being "just an API"

Most developers met Together as a place to call Llama or DeepSeek without running their own GPUs. Today the menu includes serverless inference, provisioned throughput, dedicated inference endpoints, raw GPU clusters, fine-tuning, training infrastructure, code sandboxes, and kernel and compiler work. The company says it hosts more than 200 open-source models across chat, image, audio, vision, code and embeddings.

The customer list has moved upmarket too. Together names Cognition, Decagon, ElevenLabs, Cursor and Suno among "thousands" of customers. It also cites more than 700,000 developers worldwide, without saying whether that count means registered, active or cumulative accounts. That distinction matters more than the number itself.

Pricing that is transparent and still hard to read

I would call this the menu-depth problem. Every price is public, but there are so many prices that the transparency stops helping you compare. Here is a sample from Together's pricing page:

Service Example Listed price
Serverless inference DeepSeek V4 Flash 0731 $0.14 input / $0.28 output per 1M tokens
Serverless inference Llama 3.3 70B $1.04 per 1M tokens (input and output)
Serverless inference Kimi K3 $3.00 input / $15.00 output per 1M tokens
Dedicated inference NVIDIA HGX H100 $5.49 per GPU-hour
Dedicated inference NVIDIA HGX B200 $8.99 per GPU-hour
GPU cluster H100 $1.99 preemptible, $3.99 on demand, $3.19 to $3.69 reserved
GPU cluster H200 $2.99 preemptible, $5.99 on demand
Fine-tuning (SFT) Qwen3.5 0.8B to GLM-5.1 $0.34 to $40 per 1M tokens, minimums apply
Code sandbox Compute / memory $0.0446 per vCPU-hour, $0.0149 per GiB-hour

Output tokens on Kimi K3 cost about a hundred times what DeepSeek V4 Flash output costs on the same platform. Any "Together AI is cheap" headline is meaningless until you name the model.

The rows also describe different services. An H100 at $5.49 an hour on dedicated inference comes with managed endpoints, performance guarantees and autoscaling. The same chip at $3.99 in a cluster gives you raw capacity, and your own team runs it.

Warning: Together's provisioned-throughput calculator assumes about 43,800 minutes of provisioning a month, which is every minute of every day. Its sample of 10 units shows roughly $21,600 a month and $7,236 (83%) in estimated savings, but only under the traffic assumptions it specifies. If your traffic dips overnight in IST or on weekends, you still pay for those hours.

Compute is becoming the product

The 500-megawatt line suggests where differentiation is heading. Model choice is increasingly a commodity, since the same open-weight models run on many hosts. Guaranteed access to power and chips is scarcer. That puts Together in direct competition with neoclouds and hyperscalers for the same hardware NVIDIA is pushing into the market.

The performance claims need the same caution. Together says its inference engine is "2–3 times faster than today's hyperscaler solutions." It also claims 90% faster training on HGX B200 clusters compared with previous-generation hardware, and a 24% gain in training operations from its Together Kernel Collection. Without published benchmark methodology, none of these can be checked independently. Treat them as vendor marketing until third-party tests appear.

What people are actually saying

The company's argument is about cost. Together AI states that "open-weight models can produce 6× to 20× lower costs than closed frontier models in production," and it points to Decagon, which it says cut inference costs sixfold after switching. That is one customer's result with one workload mix. It is a case study, and it does not predict what you will save.

Online discussion is more skeptical. On r/technology, users call generative AI infrastructure deeply inefficient and question whether the spending matches the value delivered. Posters on r/investing ask how fast AI hardware depreciates and what happens to capex if enterprise adoption softens. On r/ValueInvesting the recurring question is blunt: where does durable value come from if utilization falls behind capacity? On r/GenAI4all, workplace users describe employers restricting premium model access as costs rise. That is evidence that buyers are already sensitive to price.

What this means if you are buying inference

Start on serverless, measure your real token mix for at least a month, and only then consider provisioned throughput or reserved GPUs. Reserved capacity pays off only when demand is steady and near round-the-clock. With spiky traffic, per-token billing on a cheap open-weight model usually beats a reservation that sits idle.

When you compare options, count total cost of ownership:

  1. Engineering hours to run clusters yourself instead of using managed endpoints.
  2. Expected utilization, measured honestly against that 43,800-minute assumption.
  3. Latency guarantees you actually need, which justify the dedicated-inference premium.
  4. Lock-in from fine-tuned weights and reservation terms.
  5. Hardware turnover, since B200 pricing will reshape H100 economics within a single reservation period.

For broader context on which vendors earn their margins, see our piece on AI stocks facing revenue reality.

Forecast: what to watch through mid-2027

These are my expectations, not reported facts. By the first quarter of 2027, I expect Together to add more B200-class SKUs and cut H100 list rates, because the cluster table already shows H200 on-demand at $5.99. Watch for two disclosures that would change the analysis: how much of the 500 megawatts is online, and an active-developer definition to replace the 700,000 figure. If those arrive with independent benchmarks before the next funding round, the scrutiny will ease. If not, the gap between bookings and verified usage will become the central question. We will track it in The Daily Brief.

Related Reading


The Daily Brief A daily email newsletter delivering the day's trending technology, cryptocurrency, and finance news every morning.

Explore The Daily Brief

Stay ahead. For daily AI, crypto, finance & tech coverage you can trust, Veritya Daily has you covered.