Model Catalog
3 items
Applied Filters
Qwen /
Qwen3-Embedding-8B-GGUF
21
Deployed 21 times
Feature Extraction
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen3-Embedding-4B-GGUF
12
Deployed 12 times
Feature Extraction
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen3-Embedding-0.6B-GGUF
18
Deployed 18 times
Feature Extraction
Llama.cpp
Accelerated llama.cpp
F16
CPU 2x Intel Sapphire Rapids
$ 0.067