Model Catalog
9 itemsApplied Filters
Qwen /
Qwen3-Embedding-8B
Feature Extraction
vLLM
Accelerated vLLM
INF2 2x Cores
$ 1.95
Qwen /
Qwen3-Embedding-4B
Feature Extraction
vLLM
Accelerated vLLM
INF2 2x Cores
$ 1.95
Qwen /
Qwen3-Embedding-0.6B
Feature Extraction
vLLM
Accelerated vLLM
INF2 2x Cores
$ 1.95
Qwen /
Qwen3-Embedding-8B-GGUF
Feature Extraction
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen3-Embedding-4B-GGUF
Feature Extraction
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen3-Embedding-0.6B-GGUF
Feature Extraction
Llama.cpp
Accelerated llama.cpp
F16
CPU 2x Intel Sapphire Rapids
$ 0.067
Qwen /
Qwen3-Embedding-0.6B
Deployed 183 times
Feature Extraction
TEI
Accelerated Text Embeddings Inference
GPU 1x Nvidia L4
$ 0.8
Qwen /
Qwen3-Embedding-4B
Feature Extraction
TEI
Accelerated Text Embeddings Inference
GPU 1x Nvidia L4
$ 0.8
Qwen /
Qwen3-Embedding-8B
Deployed 217 times
Feature Extraction
TEI
Accelerated Text Embeddings Inference
GPU 1x Nvidia L4
$ 0.8