together.ai
The fastest cloud platform for building and running generative AI.
Together.ai Inference provides fast, scalable, and cost-efficient serverless API endpoints for deploying and fine-tuning leading open-source models like Llama-2 and Mistral. It emphasizes speed and efficiency, claiming up to 3x faster performance and 6x lower costs than competitors, alongside automatic scaling to meet growing API request volumes. The platform supports over 100 models.
Pricing: Per token usage
Resources
Together AI runs open-weight models as a hosted service behind an OpenAI-compatible API, alongside fine-tuning, dedicated endpoints, batch jobs, an evaluations API, a code-execution sandbox and custom model hosting. For teams that want the same models on reserved hardware, it also rents GPU clusters directly.
Serverless pricing is per model and per million tokens. Read on 16 August 2026 from Together's own pricing page and cross-checked against its serverless model docs: Llama 3.3 70B Instruct Turbo at $1.04 in and $1.04 out, GLM-5.2 at $1.40 and $4.40 with a 512K context window, and DeepSeek-V4-Flash at $0.14 and $0.28 with a 1M context window. Dedicated inference is listed at $5.49 per H100 GPU-hour and $8.99 for B200; GPU clusters start at $3.99 per H100-hour on demand and $3.19 reserved for 181 days or more. The model catalogue changes quickly, so treat any of these figures as needing a re-check rather than as settled.
The company was founded in 2022, is led by Vipul Ved Prakash, holds ISO 27001:2022 certification, and announced an $800 million Series C on 1 July 2026 with investors including NVIDIA, Aramco Ventures, Vista Equity and General Catalyst.
together.ai Alternatives
Explore 94 products in the Inference APIs category. View all together.ai alternatives.
LLM Tech
EU inference provider serving Qwen3.8-27B from dedicated hardware in Helsinki, with zero data retention
Varion
OpenAI-compatible proxy that trims LLM token usage before requests reach your provider
Scaleway
European serverless AI inference APIs, 100% hosted in Europe
Alibaba Cloud Model Studio
Hosted API access to Qwen and third-party models across six global regions
Compare
Also listed in
Work on together.ai? Feature it at the top of Inference APIs.
Is your product missing?