Models
Every model, one efficient API
Open-weight language, reasoning, coding, vision, embedding, and speech models, all hosted by us on 100% renewable energy and tuned to run efficiently. One catalog, one key, the lowest energy use that still does the job.
The catalog
Every model, by category
Open-weight models run on 100% renewable energy with industry-leading efficiency, reachable through one OpenAI-compatible API and a single key.
Foundation models
Open-weight foundation models we host ourselves, tuned to run at the lowest energy use for each task.
- z.ai
glm-5.3
Our flagship for agentic coding and tool use: the GLM-5.2 base with scaled post-training, a 1M-token context, and 50% higher Z.ai Code Bench scores at fewer output tokens.
- Functions
- Tool Choice
- Reasoning
- Agentic coding
- Input
- €1.10
- Cache
- €0.275
- Output
- €4.40
per 1M tokens
- Moonshot AI
kimi-k3
Open-weight 2.8T-parameter native multimodal agentic model with hybrid reasoning for long-horizon coding, knowledge work, and reasoning.
- Chat
- Hybrid reasoning
- Agentic coding
- Vision
- Input
- €3.30
- Cache
- €0.825
- Output
- €16.50
per 1M tokens
- z.ai
glm-5.3-flash
Light tier of GLM-5.3: a 320B mixture-of-experts model with 18B active, native image and video input, and the same 1M-token context at a tenth of the price.
- Functions
- Tool Choice
- Reasoning
- Vision
- Input
- €0.11
- Cache
- €0.022
- Output
- €0.44
per 1M tokens
- z.ai
glm-5.2
High-intelligence reasoning model with agentic tool use and a 1M-token context window.
- Functions
- Tool Choice
- Reasoning
- Input
- €1.10
- Cache
- €0.275
- Output
- €4.40
per 1M tokens
- DeepSeek
deepseek-v4-flash-0731
Fast, low-cost DeepSeek reasoning model with tool use and a 1M-token context.
- Functions
- Tool Choice
- Reasoning
- Input
- €0.14
- Cache
- €0.04
- Output
- €0.35
per 1M tokens
- Moonshot AI
kimi-k2.6
Agentic multimodal model with vision and a 256K-token context.
- Functions
- Tool Choice
- Reasoning
- Vision
- Input
- €0.66
- Cache
- €0.22
- Output
- €3.751
per 1M tokens
- Qwen
qwen3.5-397b-a17b
Large mixture-of-experts model for code generation and agentic tasks.
- Code
- Agentic
- Functions
- Input
- €0.70
- Output
- €4.35
per 1M tokens
- Qwen
qwen3-235b-a22b-instruct-2507
High-context multilingual reasoning across a 250K-token window.
- Multilingual
- Reasoning
- Long context
- Input
- €0.90
- Output
- €2.70
per 1M tokens
- Qwen
qwen3.6-35b-a3b
Compact frontier Qwen model for agentic and reasoning work across many languages.
- Agentic
- Reasoning
- Vision
- Code
- Input
- €0.30
- Output
- €1.80
per 1M tokens
- OpenAI
gpt-oss-120b
Open-weight 120B model with vision and long-context reasoning.
- Vision
- Reasoning
- Functions
- Input
- €0.20
- Output
- €0.70
per 1M tokens
- Mistral
mistral-small-3.2-24b-instruct-2506
Efficient instruct model with function calling and vision.
- Functions
- Vision
- Input
- €0.20
- Output
- €0.40
per 1M tokens
- Meta
llama-3.3-70b-instruct
Multilingual instruction-following at 70B for broad general use.
- Multilingual
- Input
- €1.10
- Output
- €1.10
per 1M tokens
- Mistral
mistral-medium-3.5-128b
Frontier-class reasoning, coding, and vision with a long context window.
- Reasoning
- Code
- Vision
- Input
- €1.80
- Output
- €9.00
per 1M tokens
Chat & language
GreenPT-branded open-weight models for everyday chat, reasoning, and writing, tuned for European languages.
- Google Gemma 4
gemma4
Multimodal reasoning with a long context window for documents and rich prompts.
- Vision
- Reasoning
- Long context
- Input
- €0.50
- Output
- €1.50
per 1M tokens
- GPT-OSS
green-r
Advanced reasoning, writing, and multimodal understanding with GreenPT guardrails.
- Reasoning
- Vision
- Input
- €0.35
- Output
- €0.95
per 1M tokens
- Mistral Small 3.2 24B
green-l
Fast multilingual model with Dutch grammar guardrails for European workloads.
- Multilingual
- Functions
- Input
- €0.25
- Output
- €0.80
per 1M tokens
Coding
Two models tuned specifically for software work. For a coding agent we usually recommend glm-5.3, kimi-k3, or glm-5.3-flash from the foundation models above; glm-5.2 also ships output-compression variants at the same price per token. See the coding models
- Moonshot AI
kimi-k2.7-code
Code-focused Kimi variant for software engineering, with vision and a 256K-token context.
- Functions
- Tool Choice
- Reasoning
- Vision
- Input
- €0.77
- Cache
- €0.165
- Output
- €3.85
per 1M tokens
- MiniMax
minimax-m2.5
Reasoning model tuned for agentic coding and long-horizon workflows, with a 128K max output.
- Agentic
- Reasoning
- Long-horizon
- Input
- €0.33
- Output
- €1.32
per 1M tokens
Audio & speech
Transcription and speech understanding, multilingual and accurate.
- Deepgram Nova-2
green-s
Broad language coverage for everyday workloads: twelve languages, real-time and batch, smart formatting, and speaker identification.
- Audio
- Multilingual
- Batch
- €0.12
- Live
- €0.16
per hour, 50% summer offer until 30 Sep 2026
- Deepgram Nova-3
green-s-pro
The highest accuracy available, built for medical, legal, and contact centre recordings, with a true multilingual model that follows mid-sentence language switches.
- Audio
- Multilingual
- Batch
- from €0.12
- Live
- from €0.16
per hour, 50% summer offer until 30 Sep 2026
Embeddings & retrieval
Vectors and reranking for semantic search and RAG pipelines.
- Qwen3-Embedding
green-embedding
Multilingual embeddings up to 2560 dimensions for semantic search and RAG.
- Embeddings
- Multilingual
- Price
- €0.20
per 1M tokens
- Qwen
qwen3-embedding-8b
Larger multilingual embedding model with dimensions configurable up to 4096.
- Embeddings
- Multilingual
- Price
- €0.25
per 1M tokens
- Qwen3-Reranker-4B
green-rerank
Reorders retrieved documents by true relevance, the last mile of search.
- Reranking
- Price
- €0.12
per 1M tokens
Models, in short
How do I choose a model?
Pick by capability and budget. Every model is open-weight and hosted by us, so you can match the smallest model that handles your task and get strong results at the lowest energy use and cost.
Why are these models more efficient?
They are open-weight and run on 100% renewable energy in data centres with a PUE of 1.25 and a WUE of 0.25, well below the industry averages of 1.55 and 1.8. Lighter, quantised models and automatic routing mean each request uses the least compute that still does the job.
How is pricing calculated?
Most models are priced per million input and output tokens; speech models are priced per hour of audio. Every rate, including cached input, is listed on the model cards.
See pricing per model →Is a model I need missing?
We keep adding open-weight models to the catalog. If the one you need is not here, tell us which model and what you want to use it for, and we will look at bringing it in.
Ask us for a model →How do I call a model?
Through the OpenAI-compatible API: set the base URL and key, then pass the model id. One key covers every model, plus embeddings, reranking, OCR, speech, scraping, and search.
Read the API docs →See the difference
One key for every model.
Start a free 14-day trial, no credit card. Call any model through one OpenAI-compatible API, hosted by us on 100% renewable energy and tuned for the lowest energy use from the first request.
No credit card required.
- 100% Renewable
- PUE 1.25
- Open-weight