Lección 2 · 5 min · Gratis

Gemini Flash con audio

Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

Este notebook te muestra un ejemplo de cómo usar un archivo de audio para interactuar con Gemini Flash. En este caso, usarás una grabación de sonido del discurso del Estado de la Unión del presidente John F. Kennedy en 1961.

Nota: Este notebook usa la API de Interacciones, la forma más reciente de interactuar con los modelos Gemini. ¿Buscas la versión generateContent? Consulta la rama de archivo.

Configuración

Instala las dependencias

%pip install -U -q "google-genai>=2.9.0"  # 2.0 for Interactions API

Configura tu clave de API

Para ejecutar la siguiente celda, tu clave de API debe estar almacenada en un Secreto de Colab llamado GEMINI_API_KEY. Si aún no tienes una clave de API o no sabes cómo crear un Secreto de Colab, consulta Autenticación para ver un tutorial.

from google.colab import userdata
from google import genai

GEMINI_API_KEY = userdata.get('GEMINI_API_KEY')
client = genai.Client(api_key=GEMINI_API_KEY)

Selecciona el modelo que quieres usar en esta guía:

MODEL_ID = "gemini-3.7-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.7-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input":true, isTemplate: true}

Sube un archivo de audio con la API de Archivos

Para usar un archivo de audio en tu prompt, primero debes subirlo usando la API de Archivos.

URL = "https://storage.googleapis.com/generativeai-downloads/data/State_of_the_Union_Address_30_January_1961.mp3"
!wget -q $URL -O sample.mp3
your_audio_file = client.files.upload(file='sample.mp3')

Usa el archivo en tu prompt

interaction = client.interactions.create(
    model=MODEL_ID,
    input=[
        {"type": "text", "text": "Listen carefully to the following audio file. Provide a brief summary."},
        {"type": "audio", "uri": your_audio_file.uri},
    ],
)

print(interaction.steps[-1].content[0].text)
In his first State of the Union address on January 30, 1961, President John F. Kennedy presents a candid assessment of a nation facing both "national peril and national opportunity." He details a struggling domestic economy marked by recession, high unemployment, and stagnant growth, proposing a series of legislative measures—including minimum wage increases, urban redevelopment, and expanded unemployment benefits—to stimulate recovery. 

On the international front, Kennedy addresses the heightening tensions of the Cold War, highlighting crises in Asia, Africa, and Latin America. He calls for a significant strengthening of U.S. military capabilities, specifically through increased airlift capacity and accelerated missile and submarine programs. Simultaneously, he champions a proactive foreign policy focused on international development, proposing the "Alliance for Progress" for Latin America and the creation of the Peace Corps. Kennedy concludes by emphasizing the need for scientific cooperation, arms control, and the unwavering dedication of the American people to navigate the challenging years ahead.

Audio en línea

Para solicitudes pequeñas, puedes incluir los datos de audio en la solicitud, como lo harías con las imágenes. Los datos de audio deben estar codificados en base64 para enviarse con el prompt. Usa, por ejemplo, PyDub para recortar los primeros 10 segundos del audio:

%pip install -Uq pydub
from pydub import AudioSegment
 from .audio_segment import AudioSegment
  File "/usr/local/google/home/giom/.gemini/jetski/scratch/cookbook-agentB-interactions/.venv/lib/python3.13/site-packages/pydub/audio_segment.py", line 11, in <module>
    from .utils import mediainfo_json, fsdecode
  File "/usr/local/google/home/giom/.gemini/jetski/scratch/cookbook-agentB-interactions/.venv/lib/python3.13/site-packages/pydub/utils.py", line 16, in <module>
    import pyaudioop as audioop
ModuleNotFoundError: No module named 'pyaudioop'
sound = AudioSegment.from_mp3("sample.mp3")
Traceback (most recent call last):
  File "/tmp/execute_notebook.py", line 103, in execute_notebook
    result = exec(compile(prepared, f"{name}:cell_{i}", 'exec'), ns)
  File "Audio.ipynb:cell_22", line 1, in <module>
NameError: name 'AudioSegment' is not defined
sound[:10000] # slices are in ms
Traceback (most recent call last):
  File "/tmp/execute_notebook.py", line 103, in execute_notebook
    result = exec(compile(prepared, f"{name}:cell_{i}", 'exec'), ns)
  File "Audio.ipynb:cell_23", line 1, in <module>
NameError: name 'sound' is not defined. Did you mean: 'round'?

Agrégalo a la lista de partes en el prompt:

import base64

audio_data = base64.b64encode(sound[:10000].export().read()).decode('utf-8')

interaction = client.interactions.create(
    model=MODEL_ID,
    input=[
        {"type": "text", "text": "Describe this audio clip"},
        {"type": "audio", "data": audio_data, "mime_type": "audio/mp3"},
    ],
)

print(interaction.steps[-1].content[0].text)
Traceback (most recent call last):
  File "/tmp/execute_notebook.py", line 103, in execute_notebook
    result = exec(compile(prepared, f"{name}:cell_{i}", 'exec'), ns)
  File "Audio.ipynb:cell_25", line 3, in <module>
NameError: name 'sound' is not defined. Did you mean: 'round'?

Ten en cuenta lo siguiente al proporcionar audio como datos en línea:

  • El tamaño máximo de la solicitud es de 100 MB, lo que incluye prompts de texto, instrucciones del sistema y archivos proporcionados en línea. Si el tamaño de tu archivo hace que el tamaño total de la solicitud exceda los 100 MB, entonces usa la API de Archivos para subir archivos.
  • Si estás usando una muestra de audio varias veces, es más eficiente usar la API de Archivos.

Obtén una transcripción del archivo de audio

Para obtener una transcripción, solo pídela en el prompt. Por ejemplo:

prompt = "Generate a transcript of the speech."

interaction = client.interactions.create(
    model=MODEL_ID,
    input=[
        {"type": "text", "text": prompt},
        {"type": "audio", "uri": your_audio_file.uri},
    ],
)

from IPython.display import Markdown
display(Markdown(interaction.steps[-1].content[0].text))

Haz referencia a marcas de tiempo en el archivo de audio

Un prompt puede especificar marcas de tiempo de la forma MM:SS para referirse a secciones particulares en un archivo de audio. Por ejemplo:

# Create a prompt containing timestamps.
prompt = "Provide a transcript of the speech between the timestamps 02:30 and 03:29."

interaction = client.interactions.create(
    model=MODEL_ID,
    input=[
        {"type": "text", "text": prompt},
        {"type": "audio", "uri": your_audio_file.uri},
    ],
)

display(Markdown(interaction.steps[-1].content[0].text))
Here's a transcript of the speech between the timestamps 02:30 and 03:29:

To be back among so many friends is a happy one. I am confident that that friendship will continue. Our Constitution wisely assigns both joint and separate roles to each branch of the Government; and a President and a Congress who hold each other in mutual respect will neither permit nor attempt any trespass. For my part, I shall withhold from neither the Congress nor the people any fact or report, past, present, or future, which is necessary for an informed judgment of our conduct and hazards. I shall neither shift the burden of executive decisions to the Congress, nor avoid responsibility for the outcome of those decisions.

Usa un video de YouTube

from IPython.display import display, Markdown

youtube_url = "https://www.youtube.com/watch?v=RDOMKIw1aF4" # @param {type:"string"}

prompt = """
    Analyze the following YouTube video content. Provide a concise summary covering:

    1.  **Main Thesis/Claim:** What is the central point the creator is making?
    2.  **Key Topics:** List the main subjects discussed, referencing specific examples or technologies mentioned (e.g., AI models, programming languages, projects).
    3.  **Call to Action:** Identify any explicit requests made to the viewer.
    4.  **Summary:** Provide a concise summary of the video content.

    Use the provided title, chapter timestamps/descriptions, and description text for your analysis.
"""
# Analyze the video
interaction = client.interactions.create(
    model=MODEL_ID,
    input=[
        {"type": "text", "text": prompt},
        {"type": "video", "uri": youtube_url},
    ],
)
display(Markdown(interaction.steps[-1].content[0].text))
<IPython.core.display.Markdown object>

Cuenta los tokens de audio

Puedes contar el número de tokens en tu archivo de audio usando el método count_tokens.

Los archivos de audio tienen una tasa de tokens fija por segundo (más detalles en el inicio rápido de conteo de tokens dedicado).

count_tokens_response = client.models.count_tokens(
    model=MODEL_ID,
    contents=[your_audio_file],
)

print("Audio file tokens:",count_tokens_response.total_tokens)
Audio file tokens: 83528

Próximos pasos

Referencias útiles de la API:

Más detalles sobre las capacidades de visión de la API de Gemini en la documentación.

Si quieres saber sobre la API de Archivos, consulta su referencia de la API o el inicio rápido de la API de Archivos.

Ejemplos relacionados

Consulta este ejemplo que usa archivos de audio para darte más ideas sobre lo que la API de Gemini puede hacer con ellos:

  • Comparte notas de voz con la API de Gemini y haz una lluvia de ideas.

Continúa tu descubrimiento de la API de Gemini

Echa un vistazo al inicio rápido de Audio para aprender sobre otro tipo de archivo multimedia, luego aprende más sobre cómo usar prompts con archivos multimedia en la documentación, incluyendo los formatos compatibles y la duración máxima para archivos de audio.

Lección del curso «Gemini API Cookbook (quickstarts)» de Google, publicado con licencia Apache 2.0. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Google. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios