Model Catalog

Inferentia 2
Our models optimized to run on AWS Inferentia 2, Amazon's purpose-built ML accelerator. Designed for high-throughput, cost-efficient inference at scale. Ideal for production deployments that need the performance of dedicated silicon without the cost of traditional GPU infrastructure.