Skip to content

Speech to Text

Transcribe anything. Miss nothing.

GreenPT's Speech to Text API transcribes audio to text in twelve languages, in real time or in batch. Built on Deepgram Nova-2 and Nova-3, two of the most accurate transcription models available, and run entirely on EU infrastructure so your data never leaves Europe.

  • Deepgram Nova-2 and Nova-3
  • 12 languages
  • Real-time and batch
  • GDPR-compliant
  • EU-hosted

Use cases

Wherever people talk, you need the text.

Speech to Text works everywhere audio meets workflow. Here is where our customers use it most.

  • Meeting notes

    Record once, read instantly. Turn every call into a searchable, shareable transcript without anyone lifting a pen.

  • Customer calls

    Transcribe support conversations in real time. Spot recurring issues, train your team, and keep a full searchable record of every interaction.

  • Interviews and podcasts

    Hours of audio become pages of text in seconds. Quotes, full transcripts, and chapter summaries ready to publish or analyse.

  • Accessibility

    Make every piece of audio and video content readable. Accurate captions and transcripts that help you meet legal accessibility requirements.

  • Research and analysis

    Extract insights from focus groups, user interviews, and field recordings without hours of manual transcription work.

  • Voice drafting

    Dictate instead of type. Turn spoken notes into polished emails, documents, and reports at the speed of thought.

Powered by Deepgram

The most accurate transcription models. Now on European infrastructure.

We chose Deepgram because it leads the industry on accuracy, speed, and language coverage. Green S runs Nova-2 and covers twelve languages. Green S Pro runs Nova-3, Deepgram's latest model, for the highest possible accuracy on five core languages plus multilingual mixing. Both run on our own EU-hosted servers: your audio is processed under European law and never handed off to a US provider.

Read the API docs
Models
Nova-2 / 3
Languages
12
Live latency
<1 s
Data jurisdiction
EU only

The Deepgram partnership

American technology, European control.

Deepgram builds the transcription models. We license those models and run them ourselves, on GreenPT's own servers in Europe. Deepgram is our technology supplier, not a processor of your data. So while the engine is American, your audio is only ever handled by infrastructure we operate, under European law. Here is exactly what that means for you.

Read our full privacy approach
  • Your audio never leaves the EU. Every file and live stream is processed and stored on our European servers. Nothing is sent to the United States.
  • We run the models ourselves. The models run on GreenPT's own infrastructure. There is no API call to Deepgram and no round trip to a US data centre.
  • Deepgram has no access. Deepgram cannot reach our servers, your audio, or your transcripts. They supply the technology; they never touch your data.
  • Never used for training. Your audio and transcripts are never used to train or improve any model, ours or anyone else's.
  • Outside the reach of the US CLOUD Act. Because the data and the infrastructure are fully European and Deepgram has no access, your data sits beyond US government data requests.
  • No phoning home. The models run inside our environment with no telemetry back to Deepgram. What happens in our infrastructure stays in our infrastructure.

Available models

Choose the model that fits your workload.

  • Deepgram Nova-2

    Green S

    Batch €0.52/hr · Live €0.65/hr

    Broad language coverage for everyday workloads. Twelve languages, real-time and batch, smart formatting, and speaker identification, at the lowest cost per minute.

    • 12 languages: EN, DE, ES, FR, IT, NL, PT, RO, BG, CA, DA, FI
    • Real-time and batch
    • Punctuation and smart formatting
    • Diarization V2 speaker identification
    • Sentiment, topics, and intents
  • Deepgram Nova-3

    Green S Pro

    Batch from €0.52/hr · Live from €0.78/hr

    The highest accuracy available. Built for demanding use cases: medical, legal, and contact centre recordings. Five precisely optimized languages plus a multilingual mode for conversations that mix languages.

    • 5 core languages: EN, DE, NL, SV, TR
    • Automatic language detection
    • Multilingual mode (10 languages)
    • Diarization V2 (3.3× more accurate)
    • Summarization (batch, English)

What you get

Everything your transcription pipeline needs.

  • 12 languages

    Green S covers twelve languages across Western, Northern, and Eastern Europe, including English, German, Spanish, French, Italian, Dutch, Portuguese, and more.

  • Real-time transcription

    Sub-second latency for live applications. Partial interim results let you build responsive voice UIs: live captions, call monitoring, voice assistants.

  • Any audio format

    WAV, MP3, FLAC, and all major audio formats accepted. Upload files directly or point to a URL or cloud storage bucket. Batch processing at any scale.

  • Speaker identification

    Know who said what. Powered by Deepgram Diarization V2 across all models: 3.3 times more accurate than the previous model at identifying who spoke when.

  • Sentiment, topics, and intents

    Go beyond the transcript. Detect the emotional tone of each sentence, identify key topics, and surface what speakers were trying to accomplish, without any extra code.

  • Summarization

    Get an automatic summary of long recordings alongside the full transcript. Ideal for meeting recaps, interview notes, and research. Available for batch transcription.

  • OpenAI-compatible API

    Drop into any existing integration that uses the OpenAI transcription endpoint. No code changes, no lock-in. SDKs available for Node.js and Python.

  • EU-hosted, GDPR-ready

    Your audio is processed and stored on our European servers. It is never sent to a US provider, never used to train models.

  • Sustainable infrastructure

    Every transcription runs on renewable-powered EU servers. We measure and publish the energy and CO2 footprint of every API call.

New

3.3×

more accurate speaker identification

Deepgram Diarization V2

The biggest leap in speaker identification yet.

Batch Diarization V2 uses a new speaker embedding architecture and improved clustering algorithms that cut the error rate across voice agent calls, contact centre recordings, and medical transcriptions. In head-to-head evaluations, reviewers chose V2 over the previous model more than 3.3 times as often. It works with all Nova models, supports multilingual audio, and requires no changes to your existing integration.

Read the Deepgram announcement

Energy aware

Every hour of audio has a footprint. We measure ours.

Most transcription runs on energy-hungry data centres on the other side of the world. GreenPT runs every job on European servers powered by 100% renewable energy, and we measure the energy and CO2 cost of each hour of audio you process. The greenest transcription is the one you can actually account for.

Powered by
100% renewable
Footprint
Measured per hour
Data
EU only

Pricing

Pay for the audio you transcribe.

Speech to Text is billed per hour of audio processed, with no minimums and no per-seat fees. Batch transcription is cheaper; live streaming carries a small premium for real-time delivery. Every new account starts with a 14-day free trial and full access to both models.

Start your free trial
Model Batch (per hour) Live (per hour)
Green S Nova-2 · 12 languages €0.52 €0.65
Green S Pro Nova-3 · single language €0.52 €0.78
Green S Pro Nova-3 · multilingual €0.60 €1.04

Prices exclude VAT. Speaker identification, smart formatting, sentiment, topics, and summarization are included at no extra cost. API access requires an active Pro, Teams, or API subscription.

Free for 14 days

Start transcribing today.

Full access to every language and every feature, 14 days free. No credit card required. Your audio stays in the EU from the very first second.

Cancel anytime.

  • 100% Renewable
  • EU Hosted
  • GDPR-compliant