Supported Models

evroc offers a curated catalogue of shared models across text generation, reasoning, coding, multimodal, embedding, reranking, and audio transcription. Each card below lists the model identifier, context and output limits, supported capabilities, and pricing per million tokens (input / output). Pricing is in EUR and fetched live from the evroc billing API.

GLM-5.2

zai-org/GLM-5.2

Open flagship GLM for long-horizon coding agents and million-token context work.

Context: 524k Output: 131k
Reasoning Tool calls Structured output Temperature
per 1M tokens Loading…
Gemma 4 26B A4B IT

google/gemma-4-26B-A4B-it

Open Gemma instruction model for efficient chat and self-hosted deployments.

Context: 262k Output: 33k
Reasoning Tool calls Structured output Temperature
per 1M tokens Loading…
GPT OSS 120B

openai/gpt-oss-120b

Open GPT reasoning model for self-hosted agents and controllable deployments.

Context: 66k Output: 66k
Reasoning Tool calls
per 1M tokens Loading…
Kimi K2.6

moonshotai/Kimi-K2.6

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context.

Context: 262k Output: 262k
Reasoning Tool calls Structured output Temperature
per 1M tokens Loading…
Llama 3.3 70B Instruct

nvidia/Llama-3.3-70B-Instruct-FP8

Popular open Llama workhorse for multilingual chat, coding, and self-hosting.

Context: 128k Output: 4k
Tool calls Temperature
per 1M tokens Loading…
Mistral Medium 3.5

mistralai/Mistral-Medium-3.5-128B

Balanced Mistral model for enterprise assistants, multilingual work, and tools.

Context: 262k Output: 262k
Reasoning Tool calls Structured output Temperature
per 1M tokens Loading…
Qwen3.6 35B-A3B

Qwen/Qwen3.6-35B-A3B

Open multimodal Qwen MoE for local agents that need vision, audio, and code.

Context: 262k Output: 66k
Reasoning Tool calls Structured output Temperature
per 1M tokens Loading…
Qwen3.8 27B

Qwen/Qwen3.8-27B

Dense 27B vision-language model for coding, agent tasks, and image and video understanding.

Context: 262k Output: 262k
Reasoning Tool calls Structured output Temperature
per 1M tokens Loading…
roc

evroc/roc

evroc's high-intelligence model for complex reasoning, coding, and agentic work.

Context: 262k Output: 262k
Reasoning Tool calls Structured output Temperature
per 1M tokens Loading…
E5 Multi-Lingual Large Embeddings 0.6B

intfloat/multilingual-e5-large-instruct

Multilingual embedding model for semantic search and retrieval across 100+ languages.

Context: 512 Output: 512
per 1M tokens Loading…
Qwen3 Embedding 8B

Qwen/Qwen3-Embedding-8B

High-performance text embedding model for semantic search, retrieval, and clustering.

Context: 41k Output: 4k
per 1M tokens Loading…
Qwen3 Reranker 4B

Qwen/Qwen3-Reranker-4B

Lightweight reranker for improving search relevance by reordering retrieved documents.

Context: 32k Output: 4k
per 1M tokens Loading…
Voxtral Small 24B

mistralai/Voxtral-Small-24B-2507

Multimodal Mistral model for spoken-language understanding and audio transcription.

Context: 32k Output: 32k
per 1k audio min Loading…
KB Whisper

KBLab/kb-whisper-large

Whisper-based Swedish transcription model trained by KBLab for high-quality captioning.

Context: 448 Output: 448
per 1k audio min Loading…
Whisper 3 Large

openai/whisper-large-v3

Open Whisper checkpoint for robust multilingual transcription and captioning.

Context: 448 Output: 4k
per 1k audio min Loading…
Whisper Large v3 Turbo

openai/whisper-large-v3-turbo

Speech transcription model for accurate audio-to-text and captioning workflows.

Context: 448 Output: 448
per 1k audio min Loading…

See Also