Model Catalog
6 items
Applied Filters
Qwen /
Qwen3-Embedding-8B-GGUF
14
Deployed 14 times
Feature Extraction
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen3-Embedding-4B-GGUF
3
Deployed 3 times
Feature Extraction
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen3-Embedding-0.6B-GGUF
8
Deployed 8 times
Feature Extraction
Llama.cpp
Accelerated llama.cpp
F16
CPU 2x Intel Sapphire Rapids
$ 0.067
Qwen /
Qwen3-Embedding-0.6B
138
Deployed 138 times
Feature Extraction
TEI
Accelerated Text Embeddings Inference
GPU 1x Nvidia L4
$ 0.8
Qwen /
Qwen3-Embedding-4B
99
Deployed 99 times
Feature Extraction
TEI
Accelerated Text Embeddings Inference
GPU 1x Nvidia L4
$ 0.8
Qwen /
Qwen3-Embedding-8B
166
Deployed 166 times
Feature Extraction
TEI
Accelerated Text Embeddings Inference
GPU 1x Nvidia L4
$ 0.8