Together AI
The fastest cloud platform for building and running generative AI.
Together.ai Inference provides fast, scalable, and cost-efficient serverless API endpoints for deploying and fine-tuning leading open-source models like Llama-2 and Mistral. It emphasizes speed and efficiency, claiming up to 3x faster performance and 6x lower costs than competitors, alongside automatic scaling to meet growing API request volumes. The platform supports over 100 models.
Pricing: Per token usage
Together AI prices by model
Per 1M tokens, read off Together AI's own pricing page on the date shown.
| Model | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|
| DeepSeek V4 Pro | $1.32 | $0.13 | $3.96 | 25 Sep 2026 | Listed as DeepSeek V4 Pro 0813 |
| DeepSeek V4.1 Flash | $0.30 | $0.006 | $1.20 | 25 Sep 2026 | |
| GLM 5.3 | $1.40 | $0.26 | $4.40 | 25 Sep 2026 | |
| GLM 5.3 Flash | $0.15 | $0.03 | $0.50 | 25 Sep 2026 | |
| Kimi K3 | $3.00 | $0.30 | $15.00 | 25 Sep 2026 | |
| Llama 3.3 70B Instruct | $1.04 | – | $1.04 | 25 Sep 2026 | |
| MiniMax M3 | $0.30 | $0.06 | $1.20 | 25 Sep 2026 | |
| Qwen3.8 2.4T A95B | $2.00 | $0.25 | $6.00 | 25 Sep 2026 | |
| gpt-oss-120b | $0.15 | – | $0.60 | 25 Sep 2026 |
Resources
Together AI runs open-weight models as a hosted service behind an OpenAI-compatible API, alongside fine-tuning, dedicated endpoints, batch jobs, an evaluations API, a code-execution sandbox and custom model hosting. For teams that want the same models on reserved hardware, it also rents GPU clusters directly.
Serverless pricing is per model and per million tokens. Read on 16 August 2026 from Together's own pricing page and cross-checked against its serverless model docs: Llama 3.3 70B Instruct Turbo at $1.04 in and $1.04 out, GLM-5.2 at $1.40 and $4.40 with a 512K context window, and DeepSeek-V4-Flash at $0.14 and $0.28 with a 1M context window. Dedicated inference is listed at $5.49 per H100 GPU-hour and $8.99 for B200; GPU clusters start at $3.99 per H100-hour on demand and $3.19 reserved for 181 days or more. The model catalogue changes quickly, so treat any of these figures as needing a re-check rather than as settled.
The company was founded in 2022, is led by Vipul Ved Prakash, holds ISO 27001:2022 certification, and announced an $800 million Series C on 1 July 2026 with investors including NVIDIA, Aramco Ventures, Vista Equity and General Catalyst.
Together AI Alternatives
Explore 20 products in the Fine-tuning category. View all Together AI alternatives.
OpenAI
API access to GPT, o-series reasoning, DALL-E, and Whisper models
Amazon Bedrock
Managed API access to foundation models on AWS with built-in fine-tuning and agent tooling
Replicate
Run and fine-tune open-source models. Deploy custom models at scale. All with one line of code.
Modal
Run generative AI models, large-scale batch jobs, job queues, and much more.
Klu
Collaborate on prompts, evaluate, and optimize LLM-powered Apps with Klu.
Compare
Also listed in
Work on Together AI? Feature it at the top of Fine-tuning.
Is your product missing?