Model Catalog
6 items
Applied Filters
Qwen /
Qwen3-Embedding-8B-GGUF
21
Deployed 21 times
Feature Extraction
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen3-Embedding-4B-GGUF
12
Deployed 12 times
Feature Extraction
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen3-Embedding-0.6B-GGUF
18
Deployed 18 times
Feature Extraction
Llama.cpp
Accelerated llama.cpp
F16
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen3-Embedding-0.6B
157
Deployed 157 times
Feature Extraction
TEI
Accelerated Text Embeddings Inference
GPU 1x Nvidia L4
$ 0.8
Qwen /
Qwen3-Embedding-4B
109
Deployed 109 times
Feature Extraction
TEI
Accelerated Text Embeddings Inference
GPU 1x Nvidia L4
$ 0.8
Qwen /
Qwen3-Embedding-8B
193
Deployed 193 times
Feature Extraction
TEI
Accelerated Text Embeddings Inference
GPU 1x Nvidia L4
$ 0.8