Model Catalog
11 items
Applied Filters
google /
gemma-3-12b-it
Image-Text-to-Text
TGI
Accelerated Text Generation Inference
GPU 1x Nvidia L40S
$ 1.8
google /
gemma-3-27b-it
Image-Text-to-Text
TGI
Accelerated Text Generation Inference
GPU 4x Nvidia L4
$ 3.8
ggml-org /
InternVL3-14B-Instruct-GGUF
Image-Text-to-Text
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia L4
$ 0.8
meta-llama /
Llama-3.2-11B-Vision-Instruct
Image-Text-to-Text
TGI
Accelerated Text Generation Inference
GPU 1x Nvidia L40S
$ 1.8
google /
paligemma2-10b-mix-224
Image-Text-to-Text
TGI
Accelerated Text Generation Inference
GPU 1x Nvidia L40S
$ 1.8
google /
paligemma2-10b-mix-448
Image-Text-to-Text
TGI
Accelerated Text Generation Inference
GPU 1x Nvidia L40S
$ 1.8
google /
paligemma2-3b-mix-448
Image-Text-to-Text
TGI
Accelerated Text Generation Inference
GPU 1x Nvidia L4
$ 0.8
ggml-org /
Qwen2.5-VL-3B-Instruct-GGUF
Image-Text-to-Text
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
Qwen /
Qwen2.5-VL-7B-Instruct
Image-Text-to-Text
TGI
Accelerated Text Generation Inference
GPU 1x Nvidia L40S
$ 1.8
ggml-org /
Qwen2.5-VL-7B-Instruct-GGUF
Image-Text-to-Text
Llama.cpp
Accelerated llama.cpp
Q8_0
GPU 1x Nvidia T4
$ 0.5
ggml-org /
SmolVLM2-2.2B-Instruct-GGUF
Image-Text-to-Text
Llama.cpp
Accelerated llama.cpp
F16
GPU 1x Nvidia T4
$ 0.5