Model Catalog
Llama.cpp
Discover our curated LlamaCPP model collection, optimized for fast, lightweight inference on any device. Powered by LlamaCPP’s native C++ engine, each model runs without Python dependencies, delivering low‑latency responses even on modest CPUs or GPUs.
3 items
Qwen
Qwen3-Embedding-0.6B-GGUF
/
F16
LlamaCpp Quantization
Feature Extraction
Llama.cpp Engine
Apache 2.0 License
CPU 2x Intel Sapphire Rapids
$ 0.067
ggml-org
jina-reranker-v1-turbo-en-GGUF
/
F16
LlamaCpp Quantization
Text Ranking
Llama.cpp Engine
Apache 2.0 License
CPU 8x Intel Sapphire Rapids
$ 0.268
beethogedeon
gte-Qwen2-7B-instruct-Q4_K_M-GGUF
/
Q4_K_M
LlamaCpp Quantization
Feature Extraction
Llama.cpp Engine
Apache 2.0 License
CPU 8x Intel Sapphire Rapids
$ 0.268