Telnyx offers an inference API that hosts open-weight large language models like GLM-5.2, Kimi K3, and MiniMax-M3 on globally distributed, dedicated GPU infrastructure. The platform provides OpenAI-compatible endpoints, enabling developers to switch from proprietary models and save up to 75% on token costs while maintaining sub-100 millisecond latency across multiple regions. Features include automatic scaling, in-region data privacy, function calling, structured output generation, and integrated fine-tuning, all managed through a single API key alongside Telnyx's broader communications suite.
- Models are selected for specific use cases: Kimi K3 for real-time voice, GLM-5.2 for development, and MiniMax-M3 for cost efficiency.
- Pricing starts at $0.21 per 1M tokens with no hidden GPU rental fees or compute surcharges.
- The API supports fine-tuning via the same infrastructure and requires only a base URL change for migration.
- Telnyx integrates inference with its existing voice, telephony, and storage products under one billing account.