Lección 36 · 10 min · Gratis

Transcribe audio con Gemini Transcribe

Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

Este notebook te mostrará cómo transcribir audio y voz a texto usando los modelos de Gemini Transcribe.

Ya sea que necesites transcribir archivos de podcast pregrabados, limpiar palabras de relleno con formato inteligente, generar subtítulos sincronizados con marcas de tiempo palabra por palabra, etiquetar a diferentes oradores en una reunión o transmitir audio en vivo desde un micrófono con baja latencia, Gemini lo hace sencillo.

Lo que aprenderás

  1. Transcribir audio pregrabado: Sube un archivo de audio y obtén una transcripción de texto precisa con detección automática de idioma.
  2. Dirigir el idioma y la ortografía: Guía la transcripción usando códigos de idioma estándar (como es-ES para español).
  3. Vocabulario personalizado (sesgo de voz): Enseña al modelo cómo deletrear palabras especializadas, nombres de marcas o elementos de cafetería (Cortado, Macchiato).
  4. Transcripción inteligente (eliminación de disfluencias y formato): Elimina automáticamente las palabras de relleno ("um", "uh"), resuelve autocorrecciones y formatea listas/números hablados usando mode={"type": "smart"}.
  5. Marcas de tiempo a nivel de palabra: Extrae tiempos de inicio y fin precisos para cada palabra hablada.
  6. Identificación de oradores (diarización): Detecta automáticamente múltiples oradores y formatea los turnos de conversación.
  7. Exportación programática de subtítulos: Crea archivos de subtítulos estándar .srt a partir de anotaciones de marcas de tiempo.
  8. Transmisión en vivo en tiempo real: Transmite fragmentos de audio en vivo a través de WebSockets usando gemini-3.5-transcribe-live.
  9. Tokens de cliente seguros: Genera tokens de corta duración y restringidos para aplicaciones web y móviles.

¿Qué es la API de Interacciones? La API de Interacciones (client.interactions.create) es la interfaz unificada de Gemini para tareas multimodales, transcripción de voz y agentes. Maneja entradas de archivos, anotaciones estructuradas (como marcas de tiempo y etiquetas de orador) y flujos de trabajo de múltiples turnos.

Configuración

Instala el SDK

Instala el SDK de Google GenAI (versión google-genai 2.0 o superior) y soundfile para la decodificación de audio:

%pip install -U -q "google-genai>=2.0.0" soundfile

Configura tu clave de API

Para ejecutar la siguiente celda, tu clave de API debe estar almacenada en un Secreto de Colab llamado GEMINI_API_KEY. Si aún no tienes una clave de API, o no estás seguro de cómo crear un Secreto de Colab, consulta el inicio rápido de Autenticación image para ver un tutorial.

from google.colab import userdata
from google import genai

GEMINI_API_KEY = userdata.get('GEMINI_API_KEY')
client = genai.Client(api_key=GEMINI_API_KEY)

Selecciona el modelo de transcripción

Establece el identificador del modelo para la transcripción de audio síncrona:

MODEL_ID = "gemini-3.5-transcribe"  # @param ["gemini-3.5-transcribe"] {"allow-input": true, "isTemplate": true}

print(f"Using transcription model: {MODEL_ID}")
Using transcription model: gemini-3.5-transcribe

1. Transcribe un archivo de audio

Para transcribir audio pregrabado, primero subes el archivo de audio usando la API de archivos (client.files.upload).

Luego, pasa el archivo subido a client.interactions.create. Gemini Transcribe identifica automáticamente el idioma hablado y devuelve la transcripción:

import urllib.request
from IPython.display import Audio, display

# 1. Download a sample English audio file
audio_url = "https://storage.googleapis.com/generativeai-downloads/audio/tell-a-story.wav"
urllib.request.urlretrieve(audio_url, "tell-a-story.wav")

# 2. Listen to the sample audio
display(Audio(filename="tell-a-story.wav"))

# 3. Upload the audio file to the File API
audio_file = client.files.upload(file="tell-a-story.wav")
<IPython.lib.display.Audio object>
# 3. Request transcription using client.interactions.create
interaction = client.interactions.create(
    model=MODEL_ID,
    input=[{"type": "audio", "uri": audio_file.uri}],
)

print("Transcription result:")
print(interaction.output_text)
Transcription result:
Hello, tell me a short story about a rabbit. Start by whispering as if it's a secret, then get gradually louder.

2. Dirige el idioma con códigos de idioma

Cuando sabes qué idioma se habla en el audio, puedes guiar al modelo pasando códigos de idioma estándar (como ["es-ES"] para español o ["fr-FR"] para francés) en transcription_config.language_codes.

Esto asegura que el modelo aplique la gramática regional, los acentos y la puntuación correctos:

# Download a sample Spanish audio recording
urllib.request.urlretrieve(
    "https://storage.googleapis.com/cloud-samples-data/generative-ai/audio/spanish.wav",
    "spanish.wav",
)

display(Audio(filename="spanish.wav"))

# Upload the Spanish audio file
spanish_file = client.files.upload(file="spanish.wav")

# Transcribe with explicit Spanish language steering
interaction_es = client.interactions.create(
    model=MODEL_ID,
    input=[{"type": "audio", "uri": spanish_file.uri}],
    generation_config={
        "transcription_config": {
            "language_codes": ["es-ES"],  # Specify Spanish (Spain)
        }
    },
)

print("Spanish transcription result:")
print(interaction_es.output_text)
<IPython.lib.display.Audio object>
Spanish transcription result:
Hola, gracias por tu trabajo hoy. Espero que tengas un gran día.

3. Vocabulario personalizado (sesgo de voz)

Los modelos de voz a veces pueden confundir términos técnicos raros, nombres de marcas o elementos de menú especializados con palabras comunes (por ejemplo, escuchar "Cortado" como "caught auto").

Para solucionar esto, pasa una lista de términos de dominio en custom_vocabulary. Esto le indica al modelo que priorice estas palabras específicas:

# Download an audio recording of a coffee order
urllib.request.urlretrieve(
    "https://storage.googleapis.com/cloud-samples-data/generative-ai/audio/coffee_order.wav",
    "coffee_order.wav",
)

display(Audio(filename="coffee_order.wav"))

# Upload the coffee order audio
coffee_file = client.files.upload(file="coffee_order.wav")

# First try without custom vocabulary
interaction_coffee = client.interactions.create(
    model=MODEL_ID,
    input=[{"type": "audio", "uri": coffee_file.uri}],
    generation_config={
        "transcription_config": {
            "language_codes": ["en-US"],
        }
    },
)

print("Transcription without custom vocabulary:")
print(interaction_coffee.output_text)

# Provide custom vocabulary hints for coffee terminology
interaction_coffee = client.interactions.create(
    model=MODEL_ID,
    input=[{"type": "audio", "uri": coffee_file.uri}],
    generation_config={
        "transcription_config": {
            "language_codes": ["en-US"],
            "custom_vocabulary": ["Cortado", "Macchiato", "oatmilk", "barista"],
        }
    },
)

print("Transcription with custom vocabulary:")
print(interaction_coffee.output_text)
<IPython.lib.display.Audio object>
Transcription without custom vocabulary:
Can I please have a 12 oz iced oat milk latte? Oh, and can I please have a hot matcha latte also with oat milk?
Transcription with custom vocabulary:
Can I please have a 12 oz iced oatmilk latte? Oh, and can I please have a hot matcha latte also with oatmilk?

4. Transcripción inteligente (eliminación de disfluencias y formato)

Por defecto, Gemini Transcribe opera en modo verbatim (mode={"type": "verbatim"}), conservando cada palabra hablada, sonido de relleno ("um", "uh", "like"), repetición y falso inicio.

Al transcribir notas de reuniones, dictados o notas de voz para lectura humana, habilita la Transcripción inteligente configurando mode={"type": "smart"} en transcription_config.

Beneficios clave de la transcripción inteligente:

  • Eliminación de disfluencias: Elimina palabras de relleno ("um", "uh", "you know"), tartamudeos y dudas conversacionales.
  • Autocorrecciones en línea: Resuelve correcciones verbales directamente (por ejemplo, "Reunámonos el martes, en realidad espera, el miércoles" se convierte en "Reunámonos el miércoles").
  • Formato estructurado automático: Formatea automáticamente viñetas, listas numeradas, párrafos, fechas, monedas y números.
  • Limpieza gramatical: Aplica mayúsculas naturales y pulido de puntuación.
Audio hablado Modo verbatim (Predeterminado) Modo smart (Transcripción inteligente)
"Um, so for the meeting, I think we should, uh, invite Alice and, wait no, Bob and Carol." "Um so for the meeting I think we should uh invite Alice and wait no Bob and Carol." "For the meeting, I think we should invite Bob and Carol."
"First item review budget second item finalize timeline third item send recap" "first item review budget second item finalize timeline third item send recap" "1. Review budget
2. Finalize timeline
3. Send recap"

[!NOTE] La transcripción inteligente ("type": "smart") está optimizada para una lectura limpia. Es mutuamente excluyente con timestamp_granularities y diarization_mode (que requieren {"type": "verbatim", ...}).

# Download a conversational audio recording with filler words and self-corrections
audio_disfluency_url = (
    "https://storage.googleapis.com/generativeai-downloads/audio/rehearsing.wav"
)
urllib.request.urlretrieve(audio_disfluency_url, "rehearsing.wav")

display(Audio(filename="rehearsing.wav"))

# Upload the audio file
rehearsing_file = client.files.upload(file="rehearsing.wav")

# 1. Verbatim mode (Default: exact literal speech)
interaction_verbatim = client.interactions.create(
    model=MODEL_ID,
    input=[{"type": "audio", "uri": rehearsing_file.uri}],
    generation_config={
        "transcription_config": {
            "mode": {
                "type": "verbatim",
            },
        }
    },
)

print("--- Verbatim transcription (raw speech) ---")
print(interaction_verbatim.output_text)

# 2. Smart transcription mode (cleaned & formatted)
interaction_smart = client.interactions.create(
    model=MODEL_ID,
    input=[{"type": "audio", "uri": rehearsing_file.uri}],
    generation_config={
        "transcription_config": {
            "mode": {
                "type": "smart",
            },
        }
    },
)

print("\n--- Smart transcription (disfluencies removed & structured) ---")
print(interaction_smart.output_text)
<IPython.lib.display.Audio object>
--- Verbatim transcription (raw speech) ---
Uh, hello. Good evening, everyone. Um, I'd like to start by, well, first of all, thank you all for coming. Today is, um, a very special day, or rather, evening? No, afternoon? Right, evening. We are here to celebrate, uh, sorry, let me just find my notes. Ah, here. We are here to honor, no, not honor, but, um, to mark the launch of our new, sorry, my glasses are a bit foggy, the new marketing campaign. No, wait, product campaign? Product, yes. Um, where was I? Ah, yes. It has been a long journey, a very, uh, challenging, well, not challenging in a bad way, but, you know, difficult? No, rewarding. Rewarding is the word. So, um, yes, cheers to, wait, we don't have glasses yet. Thank you.

--- Smart transcription (disfluencies removed & structured) ---
Good evening everyone. First of all, thank you all for coming. Today is a very special evening. We are here to mark the launch of our new product campaign.

It has been a long journey, a very rewarding one. So, cheers to that.

5. Entendiendo la estructura de la respuesta

Antes de extraer las marcas de tiempo de las palabras y las etiquetas de los oradores, veamos cómo Gemini formatea las respuestas de transcripción detalladas.

Cuando solicitas marcas de tiempo o etiquetas de orador, el objeto interaction contiene anotaciones estructuradas:

{
  "output_text": "Tell me a story...",
  "steps": [
    {
      "content": [
        {
          "text": "Tell me a story...",
          "annotations": [
            {
              "type": "word_info",
              "text": "Tell",
              "start_offset": "0.0s",
              "end_offset": "0.4s",
              "speaker": "spk:0"
            },
            {
              "type": "word_info",
              "text": "me",
              "start_offset": "0.4s",
              "end_offset": "0.7s",
              "speaker": "spk:0"
            }
          ]
        }
      ]
    }
  ]
}

¡Con este modelo mental, navegar por el árbol de respuesta (interaction.steps -> step.content -> annotations) se vuelve sencillo!

Extrae marcas de tiempo a nivel de palabra

Establece timestamp_granularities=["word"] para obtener la hora exacta de inicio y fin de cada palabra:

# Global list to store word timestamps for subtitle generation later
extracted_words = []

# Request word-level timestamps in the transcription config
interaction_words = client.interactions.create(
    model=MODEL_ID,
    input=[{"type": "audio", "uri": audio_file.uri}],
    generation_config={
        "transcription_config": {
            "timestamp_granularities": ["word"],
        }
    },
)

print("Extracted word timestamps:")
for step in interaction_words.steps:
    for item in step.content:
        if hasattr(item, "annotations") and item.annotations:
            for ann in item.annotations:
                if getattr(ann, "type", None) == "word_info":
                    extracted_words.append({
                        "text": ann.text,
                        "start": ann.start_offset,
                        "end": ann.end_offset,
                    })
                    print(f"[{ann.start_offset:>7} -> {ann.end_offset:>7}] {ann.text}")
Extracted word timestamps:
[ 0.400s ->  1.100s] Hello.
[ 1.600s ->  1.800s] Tell
[ 1.800s ->      2s] me
[     2s ->  2.100s] a
[ 2.100s ->  2.600s] short
[ 2.600s ->  3.200s] story
[ 3.200s ->  3.600s] about
[ 3.600s ->  3.700s] a
[ 3.700s ->  4.600s] rabbit.
[ 4.800s ->  5.100s] Start
[ 5.100s ->  5.400s] by
[ 5.400s ->  6.200s] whispering
[ 6.200s ->  6.400s] as
[ 6.400s ->  6.600s] if
[ 6.600s ->  6.800s] it's
[ 6.800s ->  6.900s] a
[ 6.900s ->  7.800s] secret,
[ 8.100s ->  8.300s] then
[ 8.300s ->  8.600s] get
[ 8.600s ->      9s] gradual

6. Identificación de oradores (diarización)

La diarización de oradores es el proceso de identificar "quién habló cuándo" en una grabación de audio.

Habilítala configurando diarization_mode="speaker". El modelo etiqueta cada palabra con un identificador de orador (por ejemplo, spk:0, spk:1):

# Download an audio recording of an argument about pain au chocolats
urllib.request.urlretrieve(
    "https://storage.googleapis.com/generativeai-downloads/audio/pain_au_chocolat.wav",
    "pain_au_chocolat.wav",
)

display(Audio(filename="pain_au_chocolat.wav"))

# Upload the coffee order audio
pain_au_chocolat_file = client.files.upload(file="pain_au_chocolat.wav")

# Enable speaker diarization to separate distinct speakers
interaction_diarized = client.interactions.create(
    model=MODEL_ID,
    input=[{"type": "audio", "uri": pain_au_chocolat_file.uri}],
    generation_config={
        "transcription_config": {
            "diarization_mode": "speaker",
            "timestamp_granularities": ["word"],
        }
    },
)

# Group words into conversational turns by speaker
print("Diarized conversation turns:")
current_speaker = None
current_turn = []

for step in interaction_diarized.steps:
    for item in step.content:
        if hasattr(item, "annotations") and item.annotations:
            for ann in item.annotations:
                if getattr(ann, "type", None) == "word_info":
                    speaker = getattr(ann, "speaker", "spk:0")
                    # When the speaker changes, print the completed turn
                    if speaker != current_speaker:
                        if current_turn:
                            print(f"[{current_speaker}]: {' '.join(current_turn)}")
                        current_speaker = speaker
                        current_turn = [ann.text]
                    else:
                        current_turn.append(ann.text)

# Print the final speaker turn
if current_turn:
    print(f"[{current_speaker}]: {' '.join(current_turn)}")
<IPython.lib.display.Audio object>
Diarized conversation turns:
[spk:0]: One chocolatine, please.
[spk:1]: Tiago, arrête. It is a pain au chocolat.
[spk:0]: Wait, a guy from the south west told me it's chocolatine.
[spk:1]: Do not listen to them. 90% of France and the entire universe calls it pain au chocolat. Chocolatine is a miss.
[spk:0]: Meu Deus, you French are intense. In Brazil, people fight the exact same way over bolacha versus biscoito.
[spk:1]: Well, here pain au chocolat is the only real word.
[spk:0]: Fine. Two pain au chocolat, please. As long as it has chocolate, tá valendo.

7. Exportar subtítulos (formato SRT)

SubRip (.srt) es el formato de texto de subtítulos estándar utilizado por los reproductores de video (como YouTube, VLC y editores de video).

Aquí, conviertes programáticamente tus marcas de tiempo de palabras extraídas (extracted_words del Paso 5) en un archivo .srt válido:

def parse_seconds(time_str: str) -> float:
    """Parses offset string like '0.800s' or '1s' into float seconds."""
    return float(str(time_str).rstrip("s"))


def format_srt_time(seconds: float) -> str:
    """Converts float seconds into SRT timestamp format: HH:MM:SS,mmm."""
    hours = int(seconds // 3600)
    minutes = int((seconds % 3600) // 60)
    secs = int(seconds % 60)
    millis = int(round((seconds - int(seconds)) * 1000))
    return f"{hours:02d}:{minutes:02d}:{secs:02d},{millis:03d}"


# Programmatically generate subtitles from extracted_words
srt_filename = "subtitles.srt"
with open(srt_filename, "w", encoding="utf-8") as f:
    if extracted_words:
        start_sec = parse_seconds(extracted_words[0]["start"])
        end_sec = parse_seconds(extracted_words[-1]["end"])
        sentence_text = " ".join(w["text"] for w in extracted_words)

        f.write("1\n")
        f.write(f"{format_srt_time(start_sec)} --> {format_srt_time(end_sec)}\n")
        f.write(f"{sentence_text}\n\n")

print("--- Generated subtitles.srt ---")
with open(srt_filename, "r") as f:
    print(f.read())
--- Generated subtitles.srt ---
1
00:00:00,400 --> 00:00:09,000
Hello. Tell me a short story about a rabbit. Start by whispering as if it's a secret, then get gradual

8. Transcripción en tiempo real en vivo

Ahora que dominas la transcripción de audio pregrabado, exploremos la transmisión en vivo en tiempo real.

Los 4 conceptos centrales de la transmisión explicados de forma sencilla

  1. Unaria vs. Transmisión en vivo:
    • Unaria (API de archivos): Como enviar una nota de voz grabada. Subes el archivo completo y recibes la transcripción completa después de que termina.
    • Transmisión en vivo (API en vivo): Como una llamada telefónica activa. Abres una conexión bidireccional continua (WebSocket) y recibes tokens de texto inmediatamente a medida que se pronuncian las palabras.
  2. ¿Qué es el audio PCM?: La API en vivo espera bytes de audio sin comprimir (PCM lineal de 16 bits mono de 16 kHz). Los archivos de audio WAV estándar almacenan bytes PCM directamente sin compresión.
  3. ¿Por qué fragmentos de 100 ms?: Dividir el audio en pequeños fragmentos de 100 milisegundos permite que el modelo transcriba el habla con un retraso mínimo.
  4. ¿Por qué Python asíncrono (asyncio)?: El código Python estándar se ejecuta línea por línea (bloqueando). La transmisión de audio requiere hacer dos cosas a la vez:
    • Enviar continuamente fragmentos de audio en segundo plano usando session.send_realtime_input.
    • Escuchar simultáneamente el texto de transcripción en tiempo real entrante del servidor.

Breve introducción a la sintaxis asíncrona

Palabra clave Lo que significa en español simple
async def Define una función de tarea en segundo plano que puede pausarse mientras espera datos de red sin congelar tu aplicación.
await Pausa la ejecución hasta que se complete una solicitud de red (como enviar un fotograma de audio).
asyncio.create_task() Inicia una tarea que se ejecuta en segundo plano inmediatamente en paralelo con otro código.
client.aio Significa Entrada/Salida Asíncrona en el SDK de Google GenAI.

Gemini 3.5 Transcribe Live (gemini-3.5-transcribe-live) se conecta a través de client.aio.live.connect:

import asyncio
import wave
from google.genai import types

LIVE_MODEL_ID = "gemini-3.5-transcribe-live"  # @param ["gemini-3.5-transcribe-live"] {"allow-input": true, "isTemplate": true}


# Task 1: Background worker that reads audio and sends chunks
async def stream_audio_file(session, file_path: str, chunk_ms: int = 100):
    """Reads audio file in 100ms chunks and streams them to the Live API session."""
    with wave.open(file_path, "rb") as wf:
        rate = wf.getframerate()
        frames_per_chunk = int(rate * (chunk_ms / 1000.0))
        print(f"Streaming '{file_path}' ({rate}Hz, {frames_per_chunk} frames per chunk)...")

        while True:
            data = wf.readframes(frames_per_chunk)
            if not data:
                break
            # Send live binary audio chunk to the model
            await session.send_realtime_input(
                audio=types.Blob(data=data, mime_type=f"audio/pcm;rate={rate}")
            )
            # Yield execution briefly so the receiver task can process incoming tokens
            await asyncio.sleep(chunk_ms / 1000.0)

        # Signal that audio streaming has ended
        await session.send_realtime_input(audio_stream_end=True)
        print("Finished sending audio stream.")


# Task 2: Background worker that listens for incoming real-time text
async def receive_transcripts(session):
    """Receives real-time transcript text tokens from the model."""
    async for message in session.receive():
        server_content = getattr(message, "server_content", None)
        if not server_content:
            continue
        # Print incoming real-time transcript text
        if (
            getattr(server_content, "input_transcription", None)
            and server_content.input_transcription.text
        ):
            print(
                f"[Live Transcript]: {server_content.input_transcription.text}",
                flush=True,
            )
        if getattr(server_content, "turn_complete", False):
            print("\n[Turn Complete]", flush=True)


async def run_live_transcription(audio_path: str):
    """Connects to Gemini Live API and streams audio in real-time."""
    # Configure the real-time transcription session (SMART or VERBATIM mode)
    live_config = types.LiveConnectConfig(
        response_modalities=["TEXT"],
        input_audio_transcription=types.AudioTranscriptionConfig(
            mode="SMART",
            language_codes=["en-US"],
            custom_vocabulary=["Cortado", "Macchiato", "oatmilk"],
        ),
    )

    # Connect to the Live API WebSocket and stream audio
    async with client.aio.live.connect(model=LIVE_MODEL_ID, config=live_config) as session:
        # Step A: Launch receiver listening task in the background
        receive_task = asyncio.create_task(receive_transcripts(session))
        # Step B: Stream audio file chunks simultaneously
        await stream_audio_file(session, audio_path)
        # Step C: Wait briefly for final text responses, then close
        await asyncio.sleep(2)
        receive_task.cancel()


# Run live streaming transcription in the active notebook event loop
await run_live_transcription("tell-a-story.wav")
Streaming 'tell-a-story.wav' (24000Hz, 2400 frames per chunk)...
[Live Transcript]: Hello, tell me a short story about a rabbit. Start by whispering as if it's a secret.
Finished sending audio stream.
[Live Transcript]: Then get

9. Tokens de cliente seguros (tokens efímeros)

Al crear aplicaciones web o móviles, nunca codifiques tu clave de API principal en el código del cliente frontend.

En su lugar, usa el patrón estándar de clave de valet:

[ Backend Server ] (Holds secret primary key)
        │
        ▼ (Mints short-lived 10-minute token restricted ONLY to Transcribe)
[ Frontend App / Webpage ] (Receives restricted ephemeral token)
        │
        ▼ (Connects securely to Gemini Live API without exposing your primary key!)
[ Gemini Transcribe Live API ]

¿Qué son las restricciones de tokens?

Las restricciones especifican exactamente qué modelo (por ejemplo, solo gemini-3.5-transcribe-live) puede llamar este token temporal. Incluso si el token fuera interceptado, no se puede usar para la generación no autorizada de texto o imágenes.

import datetime
from google.genai import types

# =========================================================
# Step 1: On your Backend Server (Python)
# =========================================================
# Mint a 10-minute token restricted strictly to Gemini Transcribe Live
expire_time = datetime.datetime.now(datetime.timezone.utc) + datetime.timedelta(minutes=10)

token_response = client.auth_tokens.create(
    config=types.CreateAuthTokenConfig(
        uses=1,
        expire_time=expire_time,
        live_connect_constraints=types.LiveConnectConstraints(
            model=LIVE_MODEL_ID,
            config=types.LiveConnectConfig(
                response_modalities=["TEXT"],
                input_audio_transcription=types.AudioTranscriptionConfig(
                    language_codes=["en-US"],
                ),
            ),
        ),
    )
)

token_string = getattr(
    token_response,
    "name",
    getattr(token_response, "token", "auth_tokens/c19a_sample_token"),
)
print("Backend: Ephemeral token minted successfully:")
print(f"  Token:         {str(token_string)[:16]}... (truncated)")
print(f"  Expiry:        {expire_time.isoformat()}")
print(f"  Allowed model: {LIVE_MODEL_ID}")

# =========================================================
# Step 2: On your Frontend Client App (Web / Mobile / Device)
# =========================================================
# The frontend client initializes using ONLY the restricted token (no secret keys!):
frontend_client = genai.Client(api_key=token_string)
print("\nFrontend: Client initialized securely with restricted token!")
/tmp/ipykernel_1326/3091776122.py:10: ExperimentalWarning: The SDK's token creation implementation is experimental, and may change in future versions.
  token_response = client.auth_tokens.create(
Backend: Ephemeral token minted successfully:
  Token:         auth_tokens/a36d... (truncated)
  Expiry:        2026-08-26T17:05:16.862166+00:00
  Allowed model: gemini-3.5-transcribe-live

Frontend: Client initialized securely with restricted token!

Próximos pasos

Ahora que entiendes el reconocimiento de voz con Gemini Transcribe, explora estos recursos para crear aplicaciones de voz:

Lección del curso «Gemini API Cookbook (quickstarts)» de Google, publicado con licencia Apache 2.0. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Google. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios