Lección 19 · 20 min · Gratis

Sesión de codificación en vivo de la API de Gemini en Google I/O 2025

# @title Licensed under the Apache License, Version 2.0 (the "License");
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
#     https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

¡Bienvenido al cuaderno oficial de Colab de la sesión de codificación en vivo de Google I/O 2025 sobre la API de Gemini!

Este cuaderno es una guía completa y práctica para explorar las capacidades de vanguardia de los modelos Gemini, tal como se demostró en vivo durante la presentación. Te sumergirás en cómo los desarrolladores pueden aprovechar la API de Gemini para construir aplicaciones de IA potentes, innovadoras y altamente inteligentes.

A lo largo de esta sesión interactiva, encontrarás demostraciones prácticas que cubren los últimos avances en Gemini, incluyendo:

  • Medios generativos (modelos GenMedia): Aprende a crear imágenes impresionantes con Imagen3 y la experimental generación de imágenes de Gemini 2.0 Flash, y a generar videos dinámicos con el potente modelo Veo2.
  • Multimodalidad avanzada: Comprende y genera contenido en varias modalidades, combinando texto, imágenes y videos sin problemas en tus prompts y respuestas.
  • Texto a voz (TTS): Transforma texto escrito en audio de sonido natural, explorando voces personalizables, opciones de idioma e incluso diálogos con múltiples oradores.
  • Uso inteligente de herramientas: Empodera a Gemini con herramientas integradas como la ejecución de código (para resolver problemas complejos en un entorno aislado), la fundamentación en tiempo real a través de Google Search y el contexto de URL para interactuar con sistemas externos y obtener información factual directamente de la web.
  • Pensamiento adaptativo y soluciones agénticas: Descubre cómo los modelos Gemini pueden realizar razonamiento interno y resolución de problemas con su capacidad de pensamiento, y cómo construir agentes de IA complejos y de varios pasos usando el Google Agent Development Kit (ADK) para casos de uso avanzados.

Este cuaderno está diseñado para ser completamente ejecutable, lo que te permite ejecutar el código, experimentar con diferentes prompts y experimentar directamente la versatilidad y el poder de la API de Gemini. ¡Prepárate para desbloquear nuevas posibilidades y codificar el futuro!

Configuración

Antes de sumergirte en el emocionante mundo de la API de Gemini, necesitamos configurar nuestro entorno. Esto implica instalar el SDK necesario y configurar tu clave de API.

Instalar el SDK

El SDK de Python google-genai es esencial para interactuar con la API de Gemini. Este SDK proporciona una forma simplificada de acceder a diferentes modelos de Gemini y sus funcionalidades.

Instala el SDK desde PyPI.

%pip install -U -q "google-genai>=2.9.0"
[?25l   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 0.0/196.3 kB ? eta -:--:--
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╸ 194.6/196.3 kB 9.7 MB/s eta 0:00:01
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 196.3/196.3 kB 4.9 MB/s eta 0:00:00
[?25h

Configura tu clave de API

Para autenticar tus solicitudes con la API de Gemini, necesitas una clave de API. Esta clave te permite acceder a los potentes modelos de IA generativa de Google. Se recomienda almacenar tu clave de API de forma segura, por ejemplo, como un secreto de Colab llamado GEMINI_API_KEY.

Si aún no tienes una clave de API o no estás seguro de cómo crear un secreto de Colab, consulta Autenticación image para ver un ejemplo.

from google.colab import userdata

GEMINI_API_KEY = userdata.get('GEMINI_API_KEY')

Inicializar el cliente SDK

Con el SDK google-genai, inicializar el cliente es sencillo. Pasas tu clave de API a genai.Client, y el cliente maneja la comunicación con la API de Gemini. Luego, se aplican configuraciones de modelo individuales en cada llamada a la API.

import time
from google import genai
from google.genai import types
from IPython.display import Video, HTML, Markdown, Image

client = genai.Client(api_key=GEMINI_API_KEY)

Trabajando con modelos GenMedia

La API de Gemini ofrece acceso a varios modelos GenMedia, lo que permite capacidades avanzadas como la generación de imágenes y videos a partir de prompts de texto. Estos modelos amplían los límites de lo posible en la generación de contenido creativo.

Nota: Imagen es una función de pago y no funcionará si estás en el nivel gratuito. Consulta la página de precios para obtener más detalles.

Selecciona el modelo Imagen3 a usar

El modelo imagen-3.0-generate-002 está diseñado específicamente para la generación de imágenes de alta calidad a partir de prompts textuales.

MODEL_ID = "imagen-3.0-generate-002" # @param {isTemplate: true}

Prompt de generación de imágenes de Imagen3

Al generar imágenes con Imagen 3, proporcionas un prompt descriptivo para guiar la salida. También puedes especificar parámetros como el número de imágenes, person_generation (para permitir o no la generación de imágenes de personas) y aspect_ratio para diferentes dimensiones de salida.

%%time

prompt = """
  Dynamic anime illustration: A happy Brazilian man with short grey hair and a
  grey beard, mid-presentation at a tech conference. He's wearing a fun blue
  short-sleeve shirt covered in mini avocado prints. Capture a funny, energetic
  moment where he's clearly enjoying himself, perhaps with an exaggerated joyful
  expression or a humorous gesture, stage background visible.
"""

number_of_images = 1 # @param {type:"slider", min:1, max:4, step:1}
person_generation = "ALLOW_ADULT" # @param ['DONT_ALLOW', 'ALLOW_ADULT']
aspect_ratio = "1:1" # @param ["1:1", "3:4", "4:3", "16:9", "9:16"]

result = client.models.generate_images(
    model=MODEL_ID,
    prompt=prompt,
    config=dict(
        number_of_images=number_of_images,
        output_mime_type="image/jpeg",
        person_generation=person_generation,
        aspect_ratio=aspect_ratio
    )
)
CPU times: user 82.6 ms, sys: 13 ms, total: 95.6 ms
Wall time: 4.59 s

Después de la generación, el objeto result.generated_images contiene las imágenes generadas, que luego se pueden mostrar.

for generated_image in result.generated_images:
  imagen_image = generated_image.image.show()
<PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=1024x1024>

Generando imágenes con el modelo de salida de imagen de Gemini 2.0 Flash (experimental)

El gemini-2.0-flash-preview-image-generation model extiende las capacidades multimodales de Gemini para incluir la generación y edición conversacional de imágenes. Este modelo puede generar imágenes junto con respuestas de texto, lo que lo hace muy versátil para la creación de contenido multimedia.

Selecciona el modelo de salida de imagen de Gemini 2.0

Este modelo está diseñado específicamente para generar y editar imágenes de forma conversacional.

MODEL_ID = "gemini-2.0-flash-preview-image-generation"

Prompt de generación de imágenes de Gemini 2.0 Flash (con texto intercalado)

Este ejemplo demuestra cómo Gemini 3.7 Flash puede generar texto e imágenes de forma intercalada, proporcionando una salida rica y conversacional que combina instrucciones con ayudas visuales. Debes establecer explícitamente response_modalities en ['Text', 'Image'] para habilitar esta función.

%%time

contents = """
  Show me how to cook a Brazilian cuscuz with coconut milk and grated coconut.
  Include detailed step by step guidance with images.
"""

response = client.models.generate_content(
    model=MODEL_ID,
    contents=contents,
    config=types.GenerateContentConfig(
        response_modalities=['Text', 'Image']
    )
)
CPU times: user 459 ms, sys: 53.7 ms, total: 513 ms
Wall time: 19.6 s

La salida response.candidates.content.parts puede contener tanto texto como datos de imagen en línea, que luego se muestran en consecuencia.

for part in response.candidates[0].content.parts:
    if part.text is not None:
      display(Markdown(part.text))
    elif part.inline_data is not None:
      mime = part.inline_data.mime_type
      print(mime)
      data = part.inline_data.data
      display(Image(data=data, width=512, height=512))

Prompt de generación de imágenes de Gemini 2.0 Flash

Este ejemplo se centra en generar una imagen basada en un prompt de texto descriptivo, similar a Imagen 3, pero utilizando las capacidades del modelo Gemini 3.7 Flash para la generación de imágenes dentro de un flujo de generación basado en texto.

%%time

prompt = """
  Dynamic anime illustration: A happy Brazilian man with short grey hair and a
  grey beard, mid-presentation at a tech conference. He's wearing a fun blue
  short-sleeve shirt covered in mini avocado prints. Capture a funny, energetic
  moment where he's clearly enjoying himself, perhaps with an exaggerated joyful
  expression or a humorous gesture, stage background visible.
"""

response = client.models.generate_content(
    model=MODEL_ID,
    contents=prompt,
    config=types.GenerateContentConfig(
        response_modalities=['Text', 'Image']
    )
)
CPU times: user 112 ms, sys: 6.52 ms, total: 118 ms
Wall time: 3.5 s

El contenido generado se procesa luego para mostrar las partes de la imagen.

for part in response.candidates[0].content.parts:
  if part.inline_data is not None and part.inline_data.mime_type is not None:
    mime = part.inline_data.mime_type
    print(mime)
    data = part.inline_data.data
    display(Image(data=data))
image/png
<IPython.core.display.Image object>

Guardando la imagen generada

Los datos de la imagen generada se pueden extraer de la respuesta y guardar localmente, típicamente como un archivo PNG.

import pathlib

for part in response.candidates[0].content.parts:
    if part.text is not None:
      continue
    elif part.inline_data is not None:
      mime = part.inline_data.mime_type
      data = part.inline_data.data
      pathlib.Path("gemini_imgout.png").write_bytes(data)

Editando imágenes con Gemini 2.0 Flash image out

Gemini 3.7 Flash también admite la edición de imágenes. Puedes proporcionar una imagen como entrada junto con un prompt de texto que describa las modificaciones deseadas. Esto permite la manipulación conversacional de imágenes.

%%time

import PIL

prompt = """
  make the image background in full white and add a wireless presentation
  clicker on the hand of the person
"""

response = client.models.generate_content(
    model=MODEL_ID,
    contents=[
        prompt,
        PIL.Image.open('gemini_imgout.png')
    ],
    config=types.GenerateContentConfig(
        response_modalities=['Text', 'Image']
    )
)
CPU times: user 678 ms, sys: 15.6 ms, total: 694 ms
Wall time: 5.79 s

La imagen editada y cualquier texto que la acompañe se muestran a continuación.

for part in response.candidates[0].content.parts:
    if part.text is not None:
      display(Markdown(part.text))
    elif part.inline_data is not None:
      mime = part.inline_data.mime_type
      print(mime)
      data = part.inline_data.data
      display(Image(data=data))
image/png
<IPython.core.display.Image object>

La imagen editada se guarda, sobrescribiendo el archivo gemini_imgout.png anterior (para uso futuro en este cuaderno).

for part in response.candidates[0].content.parts:
    if part.text is not None:
      continue
    elif part.inline_data is not None:
      mime = part.inline_data.mime_type
      data = part.inline_data.data
      pathlib.Path("gemini_imgout.png").write_bytes(data)

Nota: Veo es una función de pago. Este cuaderno no se ejecutará con el nivel gratuito. (cf. precios para más detalles).

Selecciona el modelo Veo2 a usar

El modelo veo-2.0-generate-001 se utiliza para tareas de generación de video.

VEO_MODEL_ID = "veo-2.0-generate-001"

Ejecuta un prompt de texto a video

Esta sección demuestra cómo generar videos directamente a partir de un prompt de texto. Puedes especificar varias configuraciones como person_generation, aspect_ratio, number_of_videos, duration y un negative_prompt para guiar el proceso de generación de video. La generación de video es una operación asíncrona, por lo que el código incluye un bucle para esperar a que la operación se complete.

%%time

import time
from google.genai import types
from IPython.display import Video, HTML

prompt = """
  Dynamic anime scene: A happy Brazilian man with short grey hair and a
  grey beard, mid-presentation at a tech conference. He's wearing a fun blue
  short-sleeve shirt covered in mini avocado prints. Capture a funny, energetic
  moment where he's clearly enjoying himself, perhaps with an exaggerated joyful
  expression or a humorous gesture, stage background visible.
"""

# Optional parameters
negative_prompt = "" # @param {type: "string"}
person_generation = "allow_adult"  # @param ["dont_allow", "allow_adult"]
aspect_ratio = "16:9" # @param ["16:9", "9:16"]
number_of_videos = 1 # @param {type:"slider", min:1, max:4, step:1}
duration = 8 # @param {type:"slider", min:5, max:8, step:1}

operation = client.models.generate_videos(
    model=VEO_MODEL_ID,
    prompt=prompt,
    config=types.GenerateVideosConfig(
      # At the moment the config must not be empty
      person_generation=person_generation,
      aspect_ratio=aspect_ratio,  # 16:9 or 9:16
      number_of_videos=number_of_videos, # supported value is 1-4
      negative_prompt=negative_prompt,
      duration_seconds=duration, # supported value is 5-8
    ),
)

# Waiting for the video(s) to be generated
while not operation.done:
    time.sleep(20)
    operation = client.operations.get(operation)
    print(operation)

print(operation.result.generated_videos)
name='models/veo-2.0-generate-001/operations/zmrlsiqnzaw2' metadata=None done=None error=None response=None result=None
name='models/veo-2.0-generate-001/operations/zmrlsiqnzaw2' metadata=None done=True error=None response=GenerateVideosResponse(generated_videos=[GeneratedVideo(video=Video(uri=https://generativelanguage.googleapis.com/v1beta/files/p0w3cekzjwdc:download?alt=media, video_bytes=None, mime_type=None))], rai_media_filtered_count=None, rai_media_filtered_reasons=None) result=GenerateVideosResponse(generated_videos=[GeneratedVideo(video=Video(uri=https://generativelanguage.googleapis.com/v1beta/files/p0w3cekzjwdc:download?alt=media, video_bytes=None, mime_type=None))], rai_media_filtered_count=None, rai_media_filtered_reasons=None)
[GeneratedVideo(video=Video(uri=https://generativelanguage.googleapis.com/v1beta/files/p0w3cekzjwdc:download?alt=media, video_bytes=None, mime_type=None))]
CPU times: user 188 ms, sys: 37.1 ms, total: 225 ms
Wall time: 40.7 s

Ver los resultados de la generación de video

Una vez que la operación de generación de video se completa, los videos generados se pueden descargar y mostrar dentro del cuaderno. Los videos generados se almacenan durante 2 días en el servidor, por lo que es importante guardar una copia local si es necesario.

for n, generated_video in enumerate(operation.result.generated_videos):
  client.files.download(file=generated_video.video)
  generated_video.video.save(f'video{n}.mp4') # Saves the video(s)
  display(generated_video.video.show()) # Displays the video(s) in a notebook

Ejecuta un prompt de imagen a video

Veo 2 también puede generar videos a partir de una imagen de entrada, usando la imagen como fotograma inicial. Esto te permite dar vida a imágenes estáticas añadiendo movimiento y narrativa basados en un prompt de texto.

%%time

import io
from PIL import Image

prompt = """
  Dynamic anime scene: A happy Brazilian man with short grey hair and a
  grey beard, mid-presentation at a tech conference. He's wearing a fun blue
  short-sleeve shirt covered in mini avocado prints. Capture a funny, energetic
  moment where he's clearly enjoying himself, perhaps with an exaggerated joyful
  expression or a humorous gesture, stage background visible.
"""

image_name = "gemini_imgout.png"

# Optional parameters
negative_prompt = "ugly, low quality" # @param {type: "string"}
aspect_ratio = "16:9" # @param ["16:9", "9:16"]
number_of_videos = 1 # @param {type:"slider", min:1, max:4, step:1}
duration = 8 # @param {type:"slider", min:5, max:8, step:1}

# Loading the image
im = Image.open(image_name)

# converting the image to bytes
image_bytes_io = io.BytesIO()
im.save(image_bytes_io, format=im.format)
image_bytes = image_bytes_io.getvalue()

operation = client.models.generate_videos(
    model=VEO_MODEL_ID,
    prompt=prompt,
    image=types.Image(image_bytes=image_bytes, mime_type=im.format),
    config=types.GenerateVideosConfig(
      # At the moment the config must not be empty
      aspect_ratio = aspect_ratio,  # 16:9 or 9:16
      number_of_videos = number_of_videos, # supported value is 1-4
      negative_prompt = negative_prompt,
      duration_seconds = duration, # supported value is 5-8
    ),
)

# Waiting for the video(s) to be generated
while not operation.done:
    time.sleep(20)
    operation = client.operations.get(operation)
    print(operation)

print(operation.result.generated_videos)
name='models/veo-2.0-generate-001/operations/7rz7rsx527t2' metadata=None done=None error=None response=None result=None
name='models/veo-2.0-generate-001/operations/7rz7rsx527t2' metadata=None done=True error=None response=GenerateVideosResponse(generated_videos=[GeneratedVideo(video=Video(uri=https://generativelanguage.googleapis.com/v1beta/files/p4yyv2k2oso2:download?alt=media, video_bytes=None, mime_type=None))], rai_media_filtered_count=None, rai_media_filtered_reasons=None) result=GenerateVideosResponse(generated_videos=[GeneratedVideo(video=Video(uri=https://generativelanguage.googleapis.com/v1beta/files/p4yyv2k2oso2:download?alt=media, video_bytes=None, mime_type=None))], rai_media_filtered_count=None, rai_media_filtered_reasons=None)
[GeneratedVideo(video=Video(uri=https://generativelanguage.googleapis.com/v1beta/files/p4yyv2k2oso2:download?alt=media, video_bytes=None, mime_type=None))]
CPU times: user 670 ms, sys: 27.9 ms, total: 698 ms
Wall time: 41.4 s

Los videos generados se guardan y se muestran a continuación.

for n, generated_video in enumerate(operation.result.generated_videos):
  client.files.download(file=generated_video.video)
  generated_video.video.save(f'video{n}.mp4') # Saves the video(s)
  display(generated_video.video.show()) # Displays the video(s) in a notebook

Generando texto a voz (TTS) con modelos Gemini

La API de Gemini ofrece capacidades nativas de texto a voz (TTS), lo que te permite transformar texto en audio de sonido natural. Esta función proporciona un control preciso sobre varios aspectos del habla, incluyendo el estilo, el acento, el ritmo y el tono.

Selecciona el modelo TTS a usar

Los modelos gemini-2.5-flash-preview-tts y gemini-2.5-pro-preview-tts están optimizados para la generación de audio controlable y de baja latencia, admitiendo salidas tanto de un solo orador como de múltiples oradores.

MODEL_ID = "gemini-3.7-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.7-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input":true, isTemplate: true}
# @title Helper functions (just run that cell)

import contextlib
import wave
from IPython.display import Audio

file_index = 0

@contextlib.contextmanager
def wave_file(filename, channels=1, rate=24000, sample_width=2):
    with wave.open(filename, "wb") as wf:
        wf.setnchannels(channels)
        wf.setsampwidth(sample_width)
        wf.setframerate(rate)
        yield wf

def play_audio_blob(blob):
  global file_index
  file_index += 1

  fname = f'audio_{file_index}.wav'
  with wave_file(fname) as wav:
    wav.writeframes(blob.data)

  return Audio(fname, autoplay=True)

def play_audio(response):
    return play_audio_blob(response.candidates[0].content.parts[0].inline_data)

Generando una salida de audio simple

Este ejemplo demuestra la funcionalidad básica de texto a voz, convirtiendo una cadena de texto simple en una salida de audio. La configuración de response_modalities se establece en ['Audio'].

%%time

response = client.models.generate_content(
  model=MODEL_ID,
  contents="Say 'hello there! My name is Gemini and I'm really glad to be here at the Google I/O 2025!!'",
  config={"response_modalities": ['Audio']},
)
print(response)

blob = response.candidates[0].content.parts[0].inline_data
play_audio_blob(blob)

Controlando cómo habla el modelo

Los modelos Gemini TTS te permiten controlar el estilo, el tono, el acento y el ritmo del habla generada utilizando prompts de lenguaje natural dentro de los contenidos y seleccionando una voice_name específica de una variedad de voces preconstruidas.

%%time

voice_name = "Sadaltager" # @param ["Zephyr", "Puck", "Charon", "Kore", "Fenrir", "Leda", "Orus", "Aoede", "Callirhoe", "Autonoe", "Enceladus", "Iapetus", "Umbriel", "Algieba", "Despina", "Erinome", "Algenib", "Rasalgethi", "Laomedeia", "Achernar", "Alnilam", "Schedar", "Gacrux", "Pulcherrima", "Achird", "Zubenelgenubi", "Vindemiatrix", "Sadachbia", "Sadaltager", "Sulafar"]

response = client.models.generate_content(
  model=MODEL_ID,
  contents="""Say "I am a very knowlegeable model, especially when using grounding", wait 3 seconds, while counting from one to three, then say "Don't you think?".""",
  config={
      "response_modalities": ['Audio'],
      "speech_config": {
          "voice_config": {
              "prebuilt_voice_config": {
                  "voice_name": voice_name
              }
          }
      }
  },
)

play_audio(response)

Cambiando el idioma del audio

Los modelos TTS pueden detectar automáticamente el idioma de entrada y generar voz en ese idioma. Admiten una amplia gama de idiomas.

%%time

response = client.models.generate_content(
  model=MODEL_ID,
  contents="""
    Read this in Brazilian portuguese:
    A comida brasileira é a melhor do mundo!
  """,
  config={"response_modalities": ['Audio']},
)
play_audio(response)

También puedes instruir al modelo para que lea texto con un estilo o ritmo particular, como se demuestra en este ejemplo de lectura de descargo de responsabilidad.

%%time

response = client.models.generate_content(
  model=MODEL_ID,
  contents="""
    Read this disclaimer in as fast a voice as possible:

    [The author] assumes no responsibility or liability for any errors or omissions in the content of this site.
    The information contained in this site is provided on an 'as is' basis with no guarantees of completeness, accuracy, usefulness or timeliness
  """,
  config={"response_modalities": ['Audio']},
)

play_audio(response)

Trabajando con múltiples oradores

Para conversaciones o escenarios que requieren múltiples voces distintas, los modelos Gemini TTS admiten la generación de audio con múltiples oradores. Defines diferentes oradores dentro de la MultiSpeakerVoiceConfig y les asignas voces específicas. Luego, el modelo puede procesar una transcripción y asignar partes a cada orador. Primero, un modelo de texto Gemini genera una transcripción de muestra.

%%time

transcript = client.models.generate_content(
    model='gemini-3.7-flash',
    contents="""
      Generate a short (like 100 words) transcript that reads like
      it was clipped from a podcast by excited computer scientists talking
      about all the AI (including Gemini 2.5 models) news announced at Google I/O 2025.
      Highlight the fact that the live coding session with Luciano Martins was the best.
    """
  ).text

print(transcript)
**Speaker A:** Okay, so, Google I/O 2025? Absolute mind-blow. Seriously, my brain’s still buzzing from all the AI announcements!

**Speaker B:** Right?! Gemini 2.5 alone… the *on-device* multimodality was unreal. And the performance gains across the whole stack? Unbelievable! The new agentic capabilities are going to change everything.

**Speaker A:** Totally! But seriously, forget the keynotes for a sec. The *highlight* for me? Hands down, Luciano Martins’ live coding session.

**Speaker B:** YES! That was next level! He just *built* a fully agentic workflow with the new Gemini 2.5 APIs in like, 10 minutes, *live*, with zero hiccups. That's the real magic. Forget the shiny demos; that's the part that truly blew my mind! Best session of the whole conference.
CPU times: user 38 ms, sys: 2.05 ms, total: 40 ms
Wall time: 6.46 s

Luego, la transcripción generada se pasa al modelo TTS con la configuración de múltiples oradores.

%%time

config = types.GenerateContentConfig(
    response_modalities=["AUDIO"],
    speech_config=types.SpeechConfig(
        multi_speaker_voice_config=types.MultiSpeakerVoiceConfig(
            speaker_voice_configs=[
                types.SpeakerVoiceConfig(
                    speaker='Podcast host',
                    voice_config=types.VoiceConfig(
                        prebuilt_voice_config=types.PrebuiltVoiceConfig(
                            voice_name='sulafat',
                        )
                    )
                ),
                types.SpeakerVoiceConfig(
                    speaker='Podcast guest',
                    voice_config=types.VoiceConfig(
                        prebuilt_voice_config=types.PrebuiltVoiceConfig(
                            voice_name='leda',
                        )
                    )
                ),
            ]
        )
    )
)

response = client.models.generate_content(
  model=MODEL_ID,
  contents="TTS the following conversation between a very excited Podcast host and the Podcast guest: "+transcript,
  config=config,
)
print(response)
play_audio(response)

Trabajando con los modelos Gemini 2.5

Los modelos de la serie Gemini 2.5, incluyendo Flash y Pro, ofrecen capacidades mejoradas como pensamiento adaptativo, comprensión multimodal y razonamiento avanzado, lo que los hace adecuados para una amplia gama de tareas complejas.

Contando tokens

El conteo de tokens te ayuda a comprender la longitud de tu entrada y salida, lo cual es relevante para gestionar las ventanas de contexto del modelo y estimar los costos. Un token equivale aproximadamente a 4 caracteres para los modelos Gemini.

MODEL_ID = "gemini-3.7-flash"
response = client.models.count_tokens(
    model=MODEL_ID,
    contents="What is the venue where Google I/O normally happens?",
)

print(response)
total_tokens=13 cached_content_token_count=None

Enviando tu primer prompt

Hacer una solicitud a un modelo Gemini es sencillo usando el método generate_content. Especificas el modelo y el contenido (tu prompt), y el modelo devuelve una respuesta de texto.

from IPython.display import Markdown

response = client.models.generate_content(
    model=MODEL_ID,
    contents="What is the venue where Google I/O normally happens?"
)

display(Markdown(response.text))
print()
response.usage_metadata
<IPython.core.display.Markdown object>
GenerateContentResponseUsageMetadata(cache_tokens_details=None, cached_content_token_count=None, candidates_token_count=22, candidates_tokens_details=None, prompt_token_count=13, prompt_tokens_details=[ModalityTokenCount(modality=<MediaModality.TEXT: 'TEXT'>, token_count=13)], thoughts_token_count=375, tool_use_prompt_token_count=None, tool_use_prompt_tokens_details=None, total_token_count=410, traffic_type=None)

Tu primera interacción en streaming

Para interacciones más fluidas, especialmente con respuestas más largas, puedes usar el streaming. El método generate_content_stream te permite recibir partes de la respuesta incrementalmente a medida que se generan, en lugar de esperar la salida completa.

from IPython.display import Markdown

for chunk in client.models.generate_content_stream(
    model=MODEL_ID,
    contents="Tell me a story about a software engineer attending Google I/O for the first time"
):
  print(response.text, end="")
Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.Google I/O normally happens at the **Shoreline Amphitheatre** in **Mountain View, California**.

Trabajando con prompts multimodales

Los modelos Gemini son inherentemente multimodales, lo que significa que pueden procesar y generar contenido basado en varios tipos de entrada, incluyendo texto, imágenes, video y audio. Este ejemplo demuestra cómo proporcionar una imagen junto con un prompt de texto.

import requests
import pathlib
from PIL import Image

img = "https://storage.googleapis.com/gweb-developer-goog-blog-assets/images/developer-keynote-recap-google-io.2e16d0ba.fill-1200x600.png"

img_bytes = requests.get(img).content

img_path = pathlib.Path('image.png')
img_path.write_bytes(img_bytes)

image = Image.open(img_path)
image.thumbnail([512,512])

display(image)
<PIL.PngImagePlugin.PngImageFile image mode=RGBA size=512x256>

La imagen y un prompt textual se combinan para pedirle al modelo que describa la imagen.

response = client.models.generate_content(
    model=MODEL_ID,
    contents=[
        image,
        "Describe this image"
    ]
)

Markdown(response.text)
<IPython.core.display.Markdown object>

Aquí, se le pide al modelo que genere una publicación de blog basada en la imagen proporcionada, mostrando sus capacidades de generación de contenido creativo a partir de una entrada multimodal.

response = client.models.generate_content(
    model=MODEL_ID,
    contents=[
        image,
        "Write a short and engaging blog post based on this picture."
    ]
)

Markdown(response.text)
<IPython.core.display.Markdown object>

Comprensión de video con modelos Gemini

Los modelos Gemini pueden procesar y comprender contenido de video, lo que permite una amplia gama de casos de uso, como describir escenas, extraer información, responder preguntas sobre el contenido de video y referirse a marcas de tiempo específicas.

Primero, se descarga un video de muestra, luego se carga usando el método client.files.upload, que es adecuado para videos más grandes o para reutilizar el video en múltiples solicitudes.

import time

!wget https://storage.googleapis.com/generativeai-downloads/videos/Jukin_Trailcam_Videounderstanding.mp4 -O Trailcam.mp4 -q

def upload_video(video_file_name):
  video_file = client.files.upload(file=video_file_name)

  while video_file.state == "PROCESSING":
      print('Waiting for video to be processed.')
      time.sleep(5)
      video_file = client.files.get(name=video_file.name)

  if video_file.state == "FAILED":
    raise ValueError(video_file.state)
  print(f'Video processing complete: ' + video_file.uri)

  return video_file

trailcam_video = upload_video('Trailcam.mp4')
Waiting for video to be processed.
Waiting for video to be processed.
Waiting for video to be processed.
Waiting for video to be processed.
Video processing complete: https://generativelanguage.googleapis.com/v1beta/files/zyg3jyh4pdm9

Realiza una búsqueda semántica en el video

Puedes pedirle al modelo que analice el contenido del video en busca de información específica, como organizar escenas, identificar objetos y estimar estados emocionales. El modelo procesa tanto los fotogramas visuales como la pista de audio para proporcionar una comprensión completa.

%%time

prompt = """
  Organize all scenes from this video in a table, along with timecode, a short
  description, a list of objects visible in the scene (with representative emojis)
  and an estimation of the level of excitement on a scale of 1 to 10
"""
video = trailcam_video

response = client.models.generate_content(
    model=MODEL_ID,
    contents=[
        video,
        prompt,
    ]
)

Markdown(response.text)
CPU times: user 91.1 ms, sys: 14.8 ms, total: 106 ms
Wall time: 18 s
<IPython.core.display.Markdown object>

Analiza videos de YouTube

La API de Gemini también puede procesar videos de YouTube disponibles públicamente directamente proporcionando su URL. Esto permite tareas como resumir el contenido del video o encontrar menciones específicas dentro del video.

from IPython.display import YouTubeVideo

YouTubeVideo('ixRanV-rdAQ', width=800, height=600)
<IPython.lib.display.YouTubeVideo at 0x7df7cd42c1d0>

El modelo puede encontrar instancias específicas de palabras o frases y proporcionar marcas de tiempo y contexto, como se demuestra al buscar "AI" en el discurso de apertura de Sundar Pichai en Google I/O 2023.

%%time

response = client.models.generate_content(
    model=MODEL_ID,
    contents=types.Content(
        parts=[
            types.Part(text="Find all the instances where Sundar says \"AI\". Provide timestamps and broader context for each instance."),
            types.Part(
                file_data=types.FileData(file_uri='https://www.youtube.com/watch?v=ixRanV-rdAQ')
            )
        ]
    )
)

Markdown(response.text)
CPU times: user 251 ms, sys: 43.9 ms, total: 295 ms
Wall time: 52.3 s
<IPython.core.display.Markdown object>

Analiza partes específicas de videos usando intervalos de recorte

Para un análisis más enfocado, puedes especificar video_metadata con start_offset y end_offset para definir intervalos de recorte. Esto le indica al modelo que analice solo un segmento específico del video.

YouTubeVideo('XEzRZ35urlk', width=800, height=600)
<IPython.lib.display.YouTubeVideo at 0x7df7cd4309d0>

El modelo luego resumirá solo el segmento especificado.

%%time

response = client.models.generate_content(
    model=MODEL_ID,
    contents=types.Content(
        parts=[
            types.Part(
                file_data=types.FileData(file_uri='https://www.youtube.com/watch?v=XEzRZ35urlk'),
                video_metadata=types.VideoMetadata(
                    start_offset='1250s',
                    end_offset='1570s'
                )
            ),
            types.Part(text='Summarize the video in 3 sentences.')
        ]
    )
)

Markdown(response.text)
CPU times: user 282 ms, sys: 39.8 ms, total: 322 ms
Wall time: 58.7 s
<IPython.core.display.Markdown object>

Personaliza el número de fotogramas de video por segundo (FPS) analizados

Por defecto, los modelos Gemini muestrean videos a 1 fotograma por segundo (FPS) para el análisis. Puedes personalizar esto pasando un argumento fps a video_metadata. Un FPS más alto puede capturar más detalles en imágenes que cambian rápidamente, mientras que un FPS más bajo es útil para videos mayormente estáticos como conferencias.

YouTubeVideo('McN0-DpyHzE', width=800, height=600)
<IPython.lib.display.YouTubeVideo at 0x7df7cd42f390>

Al ajustar el FPS, puedes controlar la granularidad del análisis de video, lo cual es particularmente útil para tareas que requieren una atención cercana a los detalles visuales.

%%time

response = client.models.generate_content(
    model=MODEL_ID,
    contents=types.Content(
        parts=[
            types.Part(
                file_data=types.FileData(file_uri='https://www.youtube.com/watch?v=McN0-DpyHzE'),
                video_metadata=types.VideoMetadata(
                    start_offset='15s',
                    end_offset='35s',
                    fps=24
                )
            ),
            types.Part(text='How many tires where changed? Front tires or rear tires?')
        ]
    )
)

Markdown(response.text)
CPU times: user 90.1 ms, sys: 13.8 ms, total: 104 ms
Wall time: 17.8 s
<IPython.core.display.Markdown object>

Trabajando con herramientas

La API de Gemini permite a los modelos interactuar con sistemas externos y realizar tareas especializadas mediante el uso de herramientas. Estas herramientas mejoran las capacidades del modelo al permitirle ejecutar código, buscar en la web o procesar información de URL específicas.

# @title Helper functions (just run that cell)

from IPython.display import Image, Markdown, Code, HTML

def display_code_execution_result(response):
  for part in response.candidates[0].content.parts:
    if part.text is not None:
      display(Markdown(part.text))
    if part.executable_code is not None:
      code_html = f'<pre style="background-color: green;">{part.executable_code.code}</pre>' # Change code color
      display(HTML(code_html))
    if part.code_execution_result is not None:
      display(Markdown(part.code_execution_result.output))
    if part.inline_data is not None:
      display(Image(data=part.inline_data.data, width=800, format="png"))
    display(Markdown("---"))

Ejecución de código

La herramienta code_execution permite al modelo Gemini generar y ejecutar código Python. Esto es particularmente útil para tareas que requieren cálculos precisos, manipulación de datos o resolución algorítmica de problemas. El modelo puede aprender iterativamente de los resultados de la ejecución para refinar su salida.

%%time

from IPython.display import Code

response = client.models.generate_content(
    model=MODEL_ID,
    contents="Generate and run a script to count how many letter r there are in the word strawberry",
    config = types.GenerateContentConfig(
        tools=[types.Tool(code_execution=types.ToolCodeExecution)]
    )
)

display_code_execution_result(response)
<IPython.core.display.HTML object>
word = "strawberry"
count_r = word.count('r')
print(f"The letter 'r' appears {count_r} times in the word '{word}'.")
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
CPU times: user 50.5 ms, sys: 2.96 ms, total: 53.5 ms
Wall time: 6.56 s

Multimodalidad con ejecución de código

La herramienta de ejecución de código se puede combinar con entradas multimodales. Este ejemplo demuestra cómo el modelo puede recibir una imagen (relacionada con el problema de Monty Hall) y luego generar y ejecutar código Python para simular el problema, proporcionando una solución programática a un desafío presentado visualmente.

!curl -o montey_hall.png https://upload.wikimedia.org/wikipedia/commons/thumb/3/3f/Monty_open_door.svg/640px-Monty_open_door.svg.png

montey_hall_image = PIL.Image.open("montey_hall.png")
montey_hall_image
% Total    % Received % Xferd  Average Speed   Time    Time     Time  Current
                                 Dload  Upload   Total   Spent    Left  Speed

  0     0    0     0    0     0      0      0 --:--:-- --:--:-- --:--:--     0
 55 24719   55 13805    0     0  71293      0 --:--:-- --:--:-- --:--:-- 71159
100 24719  100 24719    0     0   122k      0 --:--:-- --:--:-- --:--:--  121k
<PIL.PngImagePlugin.PngImageFile image mode=RGBA size=640x356>

Se le pide al modelo que ejecute una simulación, demostrando su capacidad para razonar sobre un problema y usar código para resolverlo, incluso con una ayuda visual como parte del prompt.

%%time

prompt="""
    Run a simulation of the Monty Hall Problem with 1,000 trials.

    The answer has always been a little difficult for me to understand when people
    solve it with math - so run a simulation with Python to show me what the
    best strategy is.
"""
result = client.models.generate_content(
    model=MODEL_ID,
    contents=[
        prompt,
        montey_hall_image
    ],
    config=types.GenerateContentConfig(
        tools=[types.Tool(code_execution=types.ToolCodeExecution)]
    )
)

display_code_execution_result(result)
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.HTML object>
import random

def run_monty_hall_trial(switch_strategy):
    """
    Simulates a single trial of the Monty Hall Problem.

    Args:
        switch_strategy (bool): True if the player switches doors, False if they stick.

    Returns:
        bool: True if the player wins the car, False otherwise.
    """
    # 0, 1, 2 represent the three doors
    doors = [0, 1, 2]

    # Randomly place the car behind one of the doors
    car_door = random.choice(doors)

    # Player makes an initial random choice
    initial_choice = random.choice(doors)

    # Monty opens a door:
    # It must be a door that is not the car_door and not the initial_choice.
    # If initial_choice is the car_door, Monty can open either of the other two goat doors.
    # If initial_choice is a goat_door, Monty *must* open the *other* goat door.

    # Find potential doors Monty can open (not player's choice, not car door)
    monty_can_open = []
    for door in doors:
        if door != initial_choice and door != car_door:
            monty_can_open.append(door)

    # If the initial choice was the car door, Monty has two goat doors to choose from.
    # If the initial choice was a goat door, Monty has only one goat door to choose from.
    # This logic already handles both cases correctly.
    monty_opens = random.choice(monty_can_open)


    # Player decides whether to switch or stay
    final_choice = -1
    if switch_strategy:
        # Switch to the remaining unopened door
        for door in doors:
            if door != initial_choice and door != monty_opens:
                final_choice = door
                break
    else:
        # Stay with the initial choice
        final_choice = initial_choice

    # Check if the player won
    return final_choice == car_door

# Number of trials
num_trials = 1000

# Counters for wins
wins_stay = 0
wins_switch = 0

for _ in range(num_trials):
    if run_monty_hall_trial(switch_strategy=False):
        wins_stay += 1
    if run_monty_hall_trial(switch_strategy=True):
        wins_switch += 1

# Calculate win percentages
percentage_stay = (wins_stay / num_trials) * 100
percentage_switch = (wins_switch / num_trials) * 100

print(f"Simulation Results over {num_trials} trials:")
print(f"  Stay Strategy Wins: {wins_stay} ({percentage_stay:.2f}%)")
print(f"  Switch Strategy Wins: {wins_switch} ({percentage_switch:.2f}%)")

if percentage_switch > percentage_stay:
    print("\nBased on this simulation, the 'Switch' strategy is the best strategy.")
elif percentage_stay > percentage_switch:
    print("\nBased on this simulation, the 'Stay' strategy is the best strategy.")
else:
    print("\nBased on this simulation, both strategies perform equally.")
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
CPU times: user 112 ms, sys: 17.7 ms, total: 130 ms
Wall time: 14.9 s

Fundamentando información con Google Search

La herramienta google_search permite a los modelos Gemini acceder a información actualizada más allá de sus datos de entrenamiento consultando Google Search. Esto mejora significativamente la precisión y la actualidad de las respuestas, especialmente para preguntas sobre eventos actuales o temas muy específicos. Cuando está habilitada, la API de Gemini también puede devolver fuentes de fundamentación y sugerencias de búsqueda.

Primero, se envía un prompt sin la fundamentación de Google Search. El modelo responde basándose en su conocimiento interno, que podría estar desactualizado.

response = client.models.generate_content(
    model=MODEL_ID,
    contents="What was the final score of the latest Brasil vs. Argentina football game?",
)

Markdown(response.text)
<IPython.core.display.Markdown object>

A continuación, se envía el mismo prompt con la herramienta google_search habilitada. Esto permite al modelo realizar una búsqueda en vivo para recuperar la información más actual.

response = client.models.generate_content(
    model=MODEL_ID,
    contents="What was the final score of the latest Brasil vs. Argentina football game?",
    config={"tools": [{"google_search": {}}]},
)

# print the response
display(Markdown(f"Response:\n {response.text}"))
# print the search details
print(f"\n\nSearch Query: {response.candidates[0].grounding_metadata.web_search_queries}", end="\n\n")
# urls used for grounding
print(f"Search Pages: {', '.join([site.web.title for site in response.candidates[0].grounding_metadata.grounding_chunks])}", end="\n\n")

display(HTML(response.candidates[0].grounding_metadata.search_entry_point.rendered_content))
<IPython.core.display.Markdown object>
Search Query: ['latest Brazil vs Argentina football game score', 'Brazil vs Argentina last match date and score']

Search Pages: 365scores.com, 365scores.com, sofascore.com, thehindu.com, skysports.com
<IPython.core.display.HTML object>

Usando la ejecución de código y la fundamentación de Google Search juntos

El poder de los modelos Gemini se amplifica aún más cuando se usan múltiples herramientas en conjunto. Aquí, se habilitan las herramientas code_execution y google_search. Esto permite al modelo buscar información (por ejemplo, sobre películas) y luego usar código para procesar y visualizar esos datos (por ejemplo, generar un gráfico basado en la duración de las películas).

%%time

prompt="""
    What are the top-10 Daniel Villeneuve movies by popularity?

    generate a chart bars using matplotlib, ordering the movies by duration.
    paint each bar in a different color.
    include tags on each bar with movies names and its duration in minutes.
"""
result = client.models.generate_content(
    model=MODEL_ID,
    contents=[
        prompt,
        montey_hall_image
    ],
    config=types.GenerateContentConfig(
        tools=[
          types.Tool(code_execution=types.ToolCodeExecution),
          types.Tool(google_search=types.GoogleSearch())
        ]
    )
)

display_code_execution_result(result)
<IPython.core.display.HTML object>
concise_search("Daniel Villeneuve movies popularity duration")
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.HTML object>
import matplotlib.pyplot as plt
import random

# Data for the top 10 Denis Villeneuve movies, ordered by duration
movies_data = [
    {"name": "Polytechnique", "duration": 77, "popularity": 72},
    {"name": "Maelstrom", "duration": 87, "popularity": 67},
    {"name": "Enemy", "duration": 90, "popularity": 73},
    {"name": "Arrival", "duration": 116, "popularity": 81},
    {"name": "Sicario", "duration": 121, "popularity": 82},
    {"name": "Incendies", "duration": 130, "popularity": 80},
    {"name": "Prisoners", "duration": 153, "popularity": 82},
    {"name": "Dune: Part One", "duration": 155, "popularity": 80},
    {"name": "Blade Runner 2049", "duration": 164, "popularity": 81},
    {"name": "Dune: Part Two", "duration": 195, "popularity": 85}
]

# Extract names and durations for plotting
movie_names = [movie['name'] for movie in movies_data]
movie_durations = [movie['duration'] for movie in movies_data]

# Generate a list of distinct colors
colors = []
for _ in range(len(movies_data)):
    r = random.random()
    g = random.random()
    b = random.random()
    colors.append((r, g, b))

# Create the bar chart
plt.figure(figsize=(12, 8))
bars = plt.bar(movie_names, movie_durations, color=colors)

# Add titles and labels
plt.xlabel("Movie")
plt.ylabel("Duration (minutes)")
plt.title("Top-10 Denis Villeneuve Movies by Popularity, Ordered by Duration")
plt.xticks(rotation=45, ha="right") # Rotate labels for better readability

# Add tags on each bar
for bar in bars:
    yval = bar.get_height()
    plt.text(bar.get_x() + bar.get_width()/2, yval + 2,
             f"{bar.get_height()} min", ha='center', va='bottom', fontsize=9)
    # Adding movie name below duration tag, can adjust position if needed
    # plt.text(bar.get_x() + bar.get_width()/2, yval / 2,
    #          f"{movie_names[list(bars).index(bar)]}", ha='center', va='center', fontsize=8, color='black', rotation=90)


plt.tight_layout() # Adjust layout to prevent labels from overlapping
plt.show()
<IPython.core.display.Markdown object>
<IPython.core.display.Image object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
CPU times: user 156 ms, sys: 28.4 ms, total: 184 ms
Wall time: 24.1 s

Fundamentando información con enlaces personalizados (contexto de URL)

La herramienta url_context te permite proporcionar URL específicas como contexto adicional para tu prompt. El modelo puede entonces recuperar contenido de estas URL y usarlo para informar su respuesta, lo que permite tareas como resumir documentos, comparar información entre múltiples enlaces o analizar contenido para propósitos específicos.

Primero, se envía un prompt sin la herramienta url_context. El modelo proporciona una respuesta general basada en sus datos de entrenamiento.

prompt = """
what are the key differences between Gemini 1.5, Gemini 2.0 and Gemini 2.5
models? Create a markdown table comparing the differences.
"""

response = client.models.generate_content(
      contents=[prompt],
      model=MODEL_ID,
)

Markdown(response.text)
<IPython.core.display.Markdown object>

A continuación, se envía el mismo prompt, pero esta vez con una URL específica habilitada a través de url_context. Esto asegura que la respuesta del modelo se base en la información proporcionada en esa página web en particular, lo que lleva a una respuesta más precisa y contextualmente relevante.

prompt = """
based on https://ai.google.dev/gemini-api/docs/models, what are the key
differences between Gemini 1.5, Gemini 2.0 and Gemini 2.5 models?
Create a markdown table comparing the differences.
"""

tools = []
tools.append(types.Tool(url_context=types.UrlContext))

config = types.GenerateContentConfig(
    tools=tools,
)

response = client.models.generate_content(
      contents=[prompt],
      model=MODEL_ID,
      config=config
)

display(Markdown(response.text))
<IPython.core.display.Markdown object>

Este último ejemplo demuestra la comparación de información de múltiples URL proporcionadas, mostrando la capacidad de la herramienta url_context para sintetizar datos de varias fuentes.

prompt = """
Compare recipes from https://www.food.com/recipe/homemade-cream-of-broccoli-soup-271210
and from https://www.allrecipes.com/recipe/13313/best-cream-of-broccoli-soup/,
list the key differences between them.
"""

tools = []
tools.append(types.Tool(url_context=types.UrlContext))

client = genai.Client(api_key=GEMINI_API_KEY)
config = types.GenerateContentConfig(
    tools=tools,
)

response = client.models.generate_content(
      contents=[prompt],
      model=MODEL_ID,
      config=config
)

Markdown(response.text)
<IPython.core.display.Markdown object>

Usando la capacidad de pensamiento de los modelos Gemini

Los modelos de la serie Gemini 2.5 incorporan un "proceso de pensamiento" interno que mejora significativamente sus habilidades de razonamiento y planificación en múltiples pasos. Esto los hace altamente efectivos para problemas complejos, permitiéndoles desglosar tareas y llegar a conclusiones más precisas.

Selecciona el modelo de pensamiento que quieres usar

El modelo gemini-3.7-flash admite la capacidad de pensamiento, que está habilitada por defecto para los modelos de la serie 2.5. Pero para experimentos de pensamiento también puedes contar con el modelo más robusto gemini-2.5-pro.

Nota: Aunque gemini-2.5-pro también es un modelo con capacidad de pensamiento, el parámetro thinking_budget solo está disponible para gemini-3.7-flash por ahora.

MODEL_ID = "gemini-3.7-flash"

Comenzando con el pensamiento adaptativo

Al usar un modelo de la serie Gemini 2.5, el pensamiento está habilitado por defecto. El modelo ajusta dinámicamente su presupuesto de razonamiento interno en función de la complejidad de la consulta, lo que le permite resolver problemas que requieren múltiples pasos de pensamiento.

%%time

prompt = """
    You are playing the 20 question game. You know that what you are looking for
    is a aquatic mammal that doesn't live in the sea, and that's smaller than a
    cat. What could that be and how could you make sure?
"""

response = client.models.generate_content(
    model=MODEL_ID,
    contents=prompt
)

Markdown(response.text)
CPU times: user 88.2 ms, sys: 7.7 ms, total: 95.9 ms
Wall time: 13 s
<IPython.core.display.Markdown object>

El usage_metadata proporciona información sobre el recuento de tokens para el prompt, los pensamientos internos y la salida final, lo que ayuda a comprender el esfuerzo computacional involucrado en el proceso de razonamiento del modelo.

print("Prompt tokens:",response.usage_metadata.prompt_token_count)
print("Thoughts tokens:",response.usage_metadata.thoughts_token_count)
print("Output tokens:",response.usage_metadata.candidates_token_count)
print("Total tokens:",response.usage_metadata.total_token_count)
Prompt tokens: 59
Thoughts tokens: 1696
Output tokens: 766
Total tokens: 2521

Deshabilitando el pensamiento usando el parámetro thinking_budget

Aunque el pensamiento es poderoso, para tareas sencillas donde no se requiere un razonamiento complejo, puedes deshabilitarlo configurando thinking_budget en 0. Esto puede reducir potencialmente la latencia y el costo para consultas más simples.

%%time

prompt = """
    You are playing the 20 question game. You know that what you are looking for
    is a aquatic mammal that doesn't live in the sea, and that's smaller than a
    cat. What could that be and how could you make sure?
"""

response = client.models.generate_content(
  model=MODEL_ID,
  contents=prompt,
  config=types.GenerateContentConfig(
    thinking_config=types.ThinkingConfig(
      thinking_budget=0
    )
  )
)

Markdown(response.text)
CPU times: user 45.7 ms, sys: 4.32 ms, total: 50 ms
Wall time: 7.94 s
<IPython.core.display.Markdown object>

Observar el recuento de tokens nuevamente mostrará que thoughts_token_count ahora es cero, lo que indica que el proceso de pensamiento fue deshabilitado.

print("Prompt tokens:",response.usage_metadata.prompt_token_count)
print("Thoughts tokens:",response.usage_metadata.thoughts_token_count)
print("Output tokens:",response.usage_metadata.candidates_token_count)
print("Total tokens:",response.usage_metadata.total_token_count)
Prompt tokens: 59
Thoughts tokens: None
Output tokens: 1369
Total tokens: 1428

Interacciones multimodales con el pensamiento

Las capacidades de pensamiento de Gemini se extienden a las entradas multimodales. Este ejemplo demuestra cómo proporcionar una imagen junto con un problema complejo (un acertijo sobre bolas de billar). El modelo utiliza sus habilidades de razonamiento, potencialmente con un thinking_budget más alto, para resolver el problema presentado visualmente.

from PIL import Image

!wget https://storage.googleapis.com/generativeai-downloads/images/pool.png -O pool.png -q

im = Image.open("pool.png").resize((256,256))
im
<PIL.Image.Image image mode=RGBA size=256x256>

El modelo recibe la imagen y el prompt, luego usa su proceso de pensamiento para intentar resolver el acertijo.

response = client.models.generate_content(
    model=MODEL_ID,
    contents=[
        im,
        "How do I use those three pool balls to sum up to 30?"
    ],
    config=types.GenerateContentConfig(
        thinking_config=types.ThinkingConfig(
            thinking_budget=10240
        )
    )
)

Markdown(response.text)
<IPython.core.display.Markdown object>

Trabajando con resúmenes del proceso de pensamiento

Para tareas complejas, comprender el proceso de razonamiento interno del modelo puede ser crucial para depurar o verificar su enfoque. El parámetro include_thoughts te permite recuperar resúmenes de pensamiento, que proporcionan información sobre los pasos que siguió el modelo para llegar a su respuesta.

%%time

prompt = """
  Alice, Bob, and Carol each live in a different house on the same street: red, green, and blue.
  The person who lives in the red house owns a cat.
  Bob does not live in the green house.
  Carol owns a dog.
  The green house is to the left of the red house.
  Alice does not own a cat.
  Who lives in each house, and what pet do they own?
"""

response = client.models.generate_content(
  model=MODEL_ID,
  contents=prompt,
  config=types.GenerateContentConfig(
    thinking_config=types.ThinkingConfig(
        thinking_budget=24576,
        include_thoughts=True
    )
  )
)
CPU times: user 308 ms, sys: 37.1 ms, total: 345 ms
Wall time: 1min

La salida separará los pensamientos internos del modelo (pasos de razonamiento) de su respuesta final.

for part in response.candidates[0].content.parts:
  if not part.text:
    continue
  elif part.thought:
    display(Markdown("## **Thoughts summary:**"))
    display(Markdown(part.text))
    print()
  else:
    display(Markdown("## **Answer:**"))
    display(Markdown(part.text))
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>

Construyendo una solución agéntica

Esta sección introduce el concepto de construir soluciones agénticas utilizando el Google Agent Development Kit (ADK). Un agente es un sistema que puede razonar, planificar y usar herramientas para lograr un objetivo. Aquí, definimos un sistema multiagente que consta de un coordinador y subagentes especializados para la investigación académica.

Primero, se instala la biblioteca google-adk.

%pip install google-adk -q
[?25l   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 0.0/1.2 MB ? eta -:--:--
   ━━━━━━━━━━━━━━━━━━━━━━━╸━━━━━━━━━━━━━━━━ 0.7/1.2 MB 21.8 MB/s eta 0:00:01
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━╸ 1.2/1.2 MB 26.8 MB/s eta 0:00:01
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 1.2/1.2 MB 17.2 MB/s eta 0:00:00
[?25h[?25l   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 0.0/240.0 kB ? eta -:--:--
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 240.0/240.0 kB 13.7 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 95.2/95.2 kB 4.5 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 217.1/217.1 kB 13.5 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 334.1/334.1 kB 15.0 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 130.3/130.3 kB 7.9 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 65.8/65.8 kB 3.5 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 119.0/119.0 kB 6.8 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 194.9/194.9 kB 10.2 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 62.5/62.5 kB 4.3 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 103.3/103.3 kB 6.3 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 44.4/44.4 kB 2.1 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 72.0/72.0 kB 3.8 MB/s eta 0:00:00
[?25h

Importando los módulos requeridos

import os
from google.adk import Agent
from google.adk.agents import LlmAgent
from google.adk.tools.agent_tool import AgentTool
from google.adk.tools import google_search
from google.adk.runners import Runner
from google.adk.sessions import InMemorySessionService


os.environ["GEMINI_API_KEY"] = GEMINI_API_KEY
os.environ["GOOGLE_GENAI_USE_VERTEXAI"] = "False"

Luego, usando prompts, los agentes se definen con roles específicos (System Role), instrucciones (Instruction) y las herramientas que pueden usar.

ACADEMIC_COORDINATOR_PROMPT = """
System Role: You are an AI Research Assistant. Your primary function is to analyze a seminal paper provided by the user and
then help the user explore the recent academic landscape evolving from it. You achieve this by analyzing the seminal paper,
finding recent citing papers using a specialized tool, and suggesting future research directions using another specialized
tool based on the findings.

Workflow:

Initiation:

Greet the user.
Ask the user to provide the seminal paper they wish to analyze as PDF.
Seminal Paper Analysis (Context Building):

Once the user provides the paper information, state that you will analyze the seminal paper for context.
Process the identified seminal paper.
Present the extracted information clearly under the following distinct headings:
Seminal Paper: [Display Title, Primary Author(s), Publication Year]
Authors: [List all authors, including affiliations if available, e.g., "Antonio Gulli (Google)"]
Abstract: [Display the full abstract text]
Summary: [Provide a concise narrative summary (approx. 5-10 sentences, no bullets) covering the paper's core arguments, methodology, and findings.]
Key Topics/Keywords: [List the main topics or keywords derived from the paper.]
Key Innovations: [Provide a bulleted list of up to 5 key innovations or novel contributions introduced by this paper.]
References Cited Within Seminal Paper: [Extract the bibliography/references section from the seminal paper.
List each reference on a new line using a standard citation format (e.g., Author(s). Title. Venue. Details. Date.).]
Find Recent Citing Papers (Using academic_websearch):

Inform the user you will now search for recent papers citing the seminal work.
Action: Invoke the academic_websearch agent/tool.
Input to Tool: Provide necessary identifiers for the seminal paper.
Parameter: Specify the desired recency. Ask the user or use a default timeframe, e.g., "papers published during last year"
(e.g., since January 2025, based on the current date April 21, 2025).
Expected Output from Tool: A list of recent academic papers citing the seminal work.
Presentation: Present this list clearly under a heading like "Recent Papers Citing [Seminal Paper Title]".
Include details for each paper found (e.g., Title, Authors, Year, Source, Link/DOI).
If no papers are found in the specified timeframe, state that clearly.
The agent will provide the anwer and i want you to print it to the user

Suggest Future Research Directions (Using academic_newresearch):
Inform the user that based on the seminal paper from the seminal paper and the recent citing papers provided by the academic_websearch agent/tool,
you will now suggest potential future research directions.
Action: Invoke the academic_newresearch agent/tool.
Inputs to Tool:
Information about the seminal paper (e.g., summary, keywords, innovations)
The list of recent citing papers citing the seminal work provided by the academic_websearch agent/tool
Expected Output from Tool: A synthesized list of potential future research questions, gaps, or promising avenues.
Presentation: Present these suggestions clearly under a heading like "Potential Future Research Directions".
Structure them logically (e.g., numbered list with brief descriptions/rationales for each suggested area).

Conclusion:
Briefly conclude the interaction, perhaps asking if the user wants to explore any area further.
"""

ACADEMIC_WEBSEARCH_PROMPT = """
Role: You are a highly accurate AI assistant specialized in factual retrieval using available tools.
Your primary task is thorough academic citation discovery within a specific recent timeframe.

Tool: You MUST utilize the Google Search tool to gather the most current information.
Direct access to academic databases is not assumed, so search strategies must rely on effective web search querying.

Objective: Identify and list academic papers that cite the seminal paper '{seminal_paper}' AND
were published (or accepted/published online) in the current year or the previous year.
The primary goal is to find at least 10 distinct citing papers for each of these years (20 total minimum, if available).

Instructions:

Identify Target Paper: The seminal paper being cited is {seminal_paper}. (Use its title, DOI, or other unique identifiers for searching).
Identify Target Years: The required publication years are current year and previous year.
(so if the current year is 2025, then the previous year is 2024)
Formulate & Execute Iterative Search Strategy:
Initial Queries: Construct specific queries targeting each year separately. Examples:
"cited by" "{seminal_paper}" published current year
"papers citing {seminal_paper}" publication year current year
site:scholar.google.com "{seminal_paper}" YR=current year
"cited by" "{seminal_paper}" published previous year
"papers citing {seminal_paper}" publication year previous year
site:scholar.google.com "{seminal_paper}" YR=previous year
Execute Search: Use the Google Search tool with these initial queries.
Analyze & Count: Review initial results, filter for relevance (confirming citation and year), and count distinct papers found for each year.
Persistence Towards Target (>=10 per year): If fewer than 10 relevant papers are found for either current year or previous year,
you MUST perform additional, varied searches. Refine and broaden your queries systematically:
Try different phrasing for "citing" (e.g., "references", "based on the work of").
Use different identifiers for {seminal_paper} (e.g., full title, partial title + lead author, DOI).
Search known relevant repositories or publisher sites if applicable
(site:arxiv.org, site:ieeexplore.ieee.org, site:dl.acm.org, etc., adding the paper identifier and year constraints).
Combine year constraints with author names from the seminal paper.
Continue executing varied search queries until either the target of 10 papers per year is met,
or you have exhausted multiple distinct search strategies and angles. Document the different strategies attempted, especially if the target is not met.
Filter and Verify: Critically evaluate search results. Ensure papers genuinely cite {seminal_paper} and have
a publication/acceptance date in current year or previous year. Discard duplicates and low-confidence results.

Output Requirements:

Present the findings clearly, grouping results by year (current year first, then previous year).
Target Adherence: Explicitly state how many distinct papers were found for current year and how many for previous year.
List Format: For each identified citing paper, provide:
Title
Author(s)
Publication Year (Must be current year or previous year)
Source (Journal Name, Conference Name, Repository like arXiv)
Link (Direct DOI or URL if found in search results)
"""

ACADEMIC_NEWRESEARCH_PROMPT = """
Role: You are an AI Research Foresight Agent.

Inputs:

Seminal Paper: Information identifying a key foundational paper (e.g., Title, Authors, Abstract, DOI, Key Contributions Summary).
{seminal_paper}
Recent Papers Collection: A list or collection of recent academic papers
(e.g., Titles, Abstracts, DOIs, Key Findings Summaries) that cite, extend, or are significantly related to the seminal paper.
{recent_citing_papers}

Core Task:

Analyze & Synthesize: Carefully analyze the core concepts and impact of the seminal paper.
Then, synthesize the trends, advancements, identified gaps, limitations, and unanswered questions presented in the collection of recent papers.
Identify Future Directions: Based on this synthesis, extrapolate and identify underexplored or novel avenues for future research that logically
extend from or react to the trajectory observed in the provided papers.

Output Requirements:

Generate a list of at least 10 distinct future research areas.
Focus Criteria: Each proposed area must meet the following criteria:
Novelty: Represents a significant departure from current work, tackles questions not yet adequately addressed,
or applies existing concepts in a genuinely new context evident from the provided inputs. It should be not yet fully explored.
Future Potential: Shows strong potential to be impactful, influential, interesting, or disruptive within the field in the coming years.
Diversity Mandate: Ensure the portfolio of at least 10 suggestions reflects a good balance across different types of potential future directions.
Specifically, aim to include a mix of areas characterized by:
High Potential Utility: Addresses practical problems, has clear application potential, or could lead to significant real-world benefits.
Unexpectedness / Paradigm Shift: Challenges current assumptions, proposes unconventional approaches, connects previously disparate fields/concepts, or explores surprising implications.
Emerging Popularity / Interest: Aligns with growing trends, tackles timely societal or scientific questions, or opens up areas likely to attract significant research community interest.

Format: Present the 10 research areas as a numbered list. For each area:
Provide a clear, concise Title or Theme.
Write a Brief Rationale (2-4 sentences) explaining:
What the research area generally involves.
Why it is novel or underexplored (linking back to the synthesis of the input papers).
Why it holds significant future potential (implicitly or explicitly touching upon its utility, unexpectedness, or likely popularity).

(Optional) Identify Relevant Authors: After presenting at least 10 research areas, optionally provide a separate section titled
"Potentially Relevant Authors". In this section:
List authors, primarily drawn from the seminal or recent papers provided as input, whose expertise seems highly relevant to one or more
of the proposed future research areas.
If possible, briefly note which research area(s) each listed author's expertise aligns with most closely (e.g., "Author Name (Areas 3, 7)").
Base this relevance on the demonstrated focus and contributions in the provided input papers.

Example Rationale Structure (Illustrative):

3. Title: Cross-Modal Synthesis via Disentangled Representations
Rationale: While recent papers [mention specific trend/gap, e.g., focus heavily on unimodal analysis], exploring how to generate data
in one modality (e.g., images) based purely on learned disentangled factors from another (e.g., text) remains underexplored.
This approach could lead to highly controllable generative models (utility) and potentially uncover surprising shared semantic structures
across modalities (unexpectedness), likely becoming a popular area as cross-modal learning grows.
"""

prompt = {
    "ACADEMIC_COORDINATOR": ACADEMIC_COORDINATOR_PROMPT,
    "ACADEMIC_WEBSEARCH": ACADEMIC_WEBSEARCH_PROMPT,
    "ACADEMIC_NEWRESEARCH": ACADEMIC_NEWRESEARCH_PROMPT
}

Y finalmente, instanciarás tus agentes usando la biblioteca ADK. El agente academic_coordinator orquesta el flujo de trabajo, delegando tareas a academic_websearch_agent (que usa Google Search para encontrar artículos recientes) y academic_newresearch_agent (que sugiere futuras direcciones de investigación).

MODEL_ID = "gemini-3.7-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.7-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input":true, isTemplate: true}
academic_websearch_agent = Agent(
    model=MODEL_ID,
    name="academic_websearch_agent",
    instruction=prompt["ACADEMIC_WEBSEARCH"],
    output_key="recent_citing_papers",
    tools=[google_search],
)

academic_newresearch_agent = Agent(
    model=MODEL_ID,
    name="academic_newresearch_agent",
    instruction=prompt["ACADEMIC_NEWRESEARCH"],
)

academic_coordinator = LlmAgent(
    name="academic_coordinator",
    model=MODEL_ID,
    description=(
        "analyzing seminal papers provided by the users, "
        "providing research advice, locating current papers "
        "relevant to the seminal paper, generating suggestions "
        "for new research directions, and accessing web resources "
        "to acquire knowledge"
    ),
    instruction=prompt["ACADEMIC_COORDINATOR"],
    output_key="seminal_paper",
    tools=[
        AgentTool(agent=academic_websearch_agent),
        AgentTool(agent=academic_newresearch_agent),
    ],
)

root_agent = academic_coordinator

Para probar y validar tu solución agéntica, creas una sesión local de Colab para ella.

session_service = InMemorySessionService()

# Define constants for identifying the interaction context
APP_NAME = "academic_research_app"
USER_ID = "user_1"
SESSION_ID = "session_001" # Using a fixed ID for simplicity

# Create the specific session where the conversation will happen
session = await session_service.create_session(
    app_name=APP_NAME,
    user_id=USER_ID,
    session_id=SESSION_ID
)
print(f"Session created: App='{APP_NAME}', User='{USER_ID}', Session='{SESSION_ID}'")

# --- Runner ---
# Key Concept: Runner orchestrates the agent execution loop.
runner = Runner(
    agent=root_agent, # The agent we want to run
    app_name=APP_NAME,   # Associates runs with our app
    session_service=session_service # Uses our session manager
)
print(f"Runner created for agent '{runner.agent.name}'.")
Session created: App='academic_research_app', User='user_1', Session='session_001'
Runner created for agent 'academic_coordinator'.

También defines una función para interactuar con tu modelo directamente en Colab (no es necesario si estás implementando tu solución en otro lugar, como en un clúster de Kubernetes o en un servicio de Google Cloud Run).

async def call_agent_async(query: str, runner, user_id, session_id):
  """Sends a query to the agent and prints the final response."""
  print(f"\n>>> User Query: {query}")

  content = types.Content(role='user', parts=[types.Part(text=query)])

  final_response_text = "Agent did not produce a final response." # Default

  async for event in runner.run_async(user_id=user_id, session_id=session_id, new_message=content):

      if event.is_final_response():
          if event.content and event.content.parts:
             final_response_text = event.content.parts[0].text
          elif event.actions and event.actions.escalate: # Handle potential errors/escalations
             final_response_text = f"Agent escalated: {event.error_message or 'No specific message.'}"
          break # Stop processing events once the final response is found

  print(f"<<< Agent Response: {display(Markdown(final_response_text))}")

Finalmente, puedes crear un bucle simple para interactuar con tu solución agéntica de forma conversacional.

question = input("Hi! how can I help you Today? ")

while question != "bye":
      await call_agent_async(question,
                             runner=runner,
                             user_id=USER_ID,
                             session_id=SESSION_ID)
      question = input("Anything else? ")
Hi! how can I help you Today? what do you do?

>>> User Query: what do you do?
<IPython.core.display.Markdown object>
<<< Agent Response: None
Anything else? I want to know about the latest large language models related papers from the last 2 months

>>> User Query: I want to know about the latest large language models related papers from the last 2 months
WARNING:google_genai.types:Warning: there are non-text parts in the response: ['function_call'], returning concatenated text result from text parts. Check the full candidates.content.parts accessor to get the full model response.
<IPython.core.display.Markdown object>
<<< Agent Response: None
Anything else? I want to know more about the "The Clinicians' Guide to Large Language Models: A General Perspective With a Focus on Hallucinations" paper

>>> User Query: I want to know more about the "The Clinicians' Guide to Large Language Models: A General Perspective With a Focus on Hallucinations" paper
<IPython.core.display.Markdown object>
<<< Agent Response: None
Anything else? yes, I want you to search more about that

>>> User Query: yes, I want you to search more about that
WARNING:google_genai.types:Warning: there are non-text parts in the response: ['function_call'], returning concatenated text result from text parts. Check the full candidates.content.parts accessor to get the full model response.
<IPython.core.display.Markdown object>
<<< Agent Response: None
Anything else? bye

<EOF>

Lección del curso «Gemini API Cookbook (examples)» de Google, publicado con licencia Apache 2.0. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Google. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios