Skip to content

Available Models

uniGPT offers a variety of models — both internal (hosted on-premises at the university) and external (provided via cloud partners).

We always support the latest model generation as well as latest-1 (the previous generation), so you can keep using familiar models while new ones are introduced.

For balance reset details and practical advice on conserving credits, see Usage Limits & Balance.

Choosing the Right Model

Rule of thumb: start with a cheap model and only switch to a more expensive one if the result isn't good enough.

  • Everyday tasks (default): DeepSeek V4 Flash for most requests (possible bias on China-related political topics). gemma-4-31B-it is a strong alternative, especially for writing.
  • Up-to-date or specific information from the web: Gemini Flash 3.8 (stronger) or GPT 5.6 Luna (cheaper) with web search enabled.
  • Harder tasks (e.g. programming, complex analysis): GPT 5.6 Sol
  • Only for really hard tasks: GPT-6 Astra, Claude Opus 5. Both are extremely capable but by far the most expensive options per task.
  • Medical text & image tasks: medgemma-1.5-4b-it
  • Sensitive/confidential data: Always use internal models, regardless of task difficulty. For coding, use Qwen3.8-27B; for demanding tasks, gpt-oss-120b. Do not enable web search.

Internal Models (On-Premises)

These models run entirely on university infrastructure. Your data never leaves the university, making them fully GDPR-compliant and suitable for sensitive data.

Costs for internal models are given as credit costs per 1M tokens (input / output).

Model Description Input / output cost (credits)
DeepSeek V4 Flash Fast, cost-effective reasoning model; multimodal (Vision) with image support 0.14 / 0.28
Qwen3.8-27B Fast, capable general-purpose model 2 / 6
Qwen3.5-35B-A3B Chinese advanced reasoning and coding model with image support 0.035 / 0.138
Muse-Glimmer-30B Open-weight model by Meta, successor to Llama-3.3
gemma-4-31B-it Efficient multimodal model with image support 0.09 / 0.16
medgemma-1.5-4b-it Specialized multimodal model for medical text and image understanding 0.02 / 0.04
gpt-oss-120b Large internal model for complex tasks 0.15 / 0.6
mistral-small-4 Hybrid model unifying instruct, reasoning, and coding capabilities 0.15 / 0.2

Retired models

Gemma 3 and Mistral Small (Legacy) have been retired. Llama-3.3-70B has been replaced by Muse-Glimmer-30B.

Best for sensitive data

Use internal models when working with personal data, unpublished research, or any data subject to GDPR restrictions.

External Models (Cloud Providers)

These models are accessed via cloud providers through the GÉANT OCRE framework. While they offer state-of-the-art performance, data is processed externally.

Costs are given as credit costs per 1M tokens (input / output).

Model Provider Notes Input / output cost (credits)
GPT-6 Astra OpenAI New flagship model, available for employees only 10 / 50
GPT 5.6 Sol OpenAI Flagship reasoning model 5 / 30
GPT 5.6 Luna OpenAI Cost-efficient GPT-5.6 variant 1 / 6
Claude Opus 5 Anthropic (via Google Cloud) Strongest Claude model for deep analysis 5 / 25
Claude Sonnet 5 Anthropic (via Google Cloud) Balanced writing and analysis model 2 / 10
Claude Haiku 4.5 Anthropic (via Google Cloud) Fastest and cheapest Claude option 1 / 5
Gemini Flash 3.8 Google Cloud Latest fast multimodal model, available for all 0.75 / 3.75
Gemini 3.1 Pro Google Cloud Older and larger multimodal model 2 / 12

Costs

Credit costs are per 1M tokens (input / output) and equal the providers' base list prices. uniGPT may still apply its own markup, exchange rate, caching rules, or internal-model pricing.

Cost awareness

External models cost significantly more per token than internal models. The university receives only a minimal discount (<5%) from providers. Please use external models sparingly and prefer internal models or cost-effective options like DeepSeek V4 Flash, GPT 5.6 Luna, or Gemini Flash 3.8 when possible. Note in particular that GPT-6 Astra and Opus are extremely capable but expensive — reserve them for complex tasks and use cheaper models for standard requests or testing. Otherwise, user limits may need to be reduced.

For current balance behavior and reset timing, see Usage Limits & Balance.