Model Catalog
Inferentia 2
Our models optimized to run on AWS Inferentia 2, Amazon's purpose-built ML accelerator. Designed for high-throughput, cost-efficient inference at scale. Ideal for production deployments that need the performance of dedicated silicon without the cost of traditional GPU infrastructure.
7 items
Qwen
Qwen3-Embedding-0.6B
Deployed 198 times
Feature Extraction
TEI Engine
Apache 2.0 License
GPU 1x Nvidia L4
$ 0.8
Qwen
Qwen3-Embedding-4B
Feature Extraction
TEI Engine
Apache 2.0 License
GPU 1x Nvidia L4
$ 0.8
Qwen
Qwen3-Embedding-8B
Deployed 348 times
Feature Extraction
TEI Engine
Apache 2.0 License
GPU 1x Nvidia L4
$ 0.8
meta-llama
Meta-Llama-3-8B-Instruct
Text Generation
vLLM Engine
Llama 3 License
GPU 1x Nvidia A100
$ 2.5
meta-llama
Llama-3.1-8B-Instruct
Deployed 557 times
Text Generation
vLLM Engine
Llama 3.1 License
GPU 1x Nvidia A100
$ 2.5
meta-llama
Llama-3.2-1B-Instruct
Text Generation
vLLM Engine
Llama 3.2 License
GPU 1x Nvidia L4
$ 0.8
meta-llama
Llama-3.2-3B-Instruct
Deployed 159 times
Text Generation
vLLM Engine
Llama 3.2 License
GPU 1x Nvidia L40S
$ 1.8