Introducción a la API de Gemini
Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
🔄 Nuevo: API de Interacciones
La API de Interacciones (
client.interactions.create) es ahora la forma recomendada de usar la API de Gemini. Ofrece:
- Estado de conversación en el servidor — no necesitas gestionar el historial de chat manualmente, solo pasa
previous_interaction_id- Entrada multimodal unificada — texto, imágenes, audio, video y documentos en una sola lista
- Orquestación de herramientas integrada — Google Search, ejecución de código y llamadas a funciones con definiciones de herramientas simplificadas
- Streaming mediante eventos — control granular con eventos tipados (
ContentDelta,ContentStart, etc.)👉 Empieza con el nuevo notebook de Introducción que usa la API de Interacciones.
Este notebook usa la API heredada
generateContenty se mantiene como referencia.
Gemini 3 Pro/Flash: Si solo te interesan las nuevas capacidades de los modelos Gemini 3 (niveles de pensamiento, resolución de medios y firmas de pensamiento), salta directamente a la sección dedicada al final de este notebook.
El SDK de Google Gen AI proporciona una interfaz unificada a los modelos Gemini a través de la API de Desarrolladores de Gemini y la API de Gemini en Vertex AI. Con algunas excepciones, el código que se ejecuta en una plataforma se ejecutará en ambas. Este notebook usa la API de Desarrolladores.
Este notebook te guiará a través de:
- Instalar y configurar el SDK de Google GenAI
- Prompting de texto y multimodal
- Configurar instrucciones del sistema
- Controlar el proceso de pensamiento
- Contar tokens
- Configurar filtros de seguridad
- Iniciar un chat de varias interacciones
- Generar un flujo de contenido y enviar solicitudes asíncronas
- Controlar la salida generada
- Usar llamadas a funciones
- Fundamentar tus solicitudes usando cargas de archivos, Google Search, Google Maps, Youtube o añadiendo URLs a tu prompt
- Usar almacenamiento en caché de contexto
- Generar embeddings
Más detalles sobre el SDK en la documentación.
Los modelos específicos de funciones tienen sus propias guías dedicadas:
- Generación de podcasts y voz usando Gemini TTS
, - Interacción en vivo con Gemini Live
, - Generación de imágenes usando Imagen
, - Generación de video usando Veo
, - Generación de música usando Lyria RealTime
.
Configuración
Instalar el SDK
Instala el SDK desde PyPI. Se recomienda usar siempre la última versión.
%pip install -U -q "google-genai>=2.9.0" # 2.0 for Interactions API
[1m[[0m[34;49mnotice[0m[1;39;49m][0m[39;49m A new release of pip is available: [0m[31;49m23.2.1[0m[39;49m -> [0m[32;49m26.1.2[0m
[1m[[0m[34;49mnotice[0m[1;39;49m][0m[39;49m To update, run: [0m[32;49mpip3 install --upgrade pip[0m
Note: you may need to restart the kernel to use updated packages.
Configurar tu clave de API
Para ejecutar la siguiente celda, tu clave de API debe estar almacenada en un Secreto de Colab llamado GEMINI_API_KEY. Si aún no tienes una clave de API o no estás seguro de cómo crear un Secreto de Colab, consulta Autenticación
para ver un tutorial.
import os
GEMINI_API_KEY = os.environ.get("GEMINI_API_KEY", "")
Inicializar el cliente del SDK
Con el nuevo SDK, ahora solo necesitas inicializar un cliente con tu clave de API (o OAuth si usas Vertex AI). El modelo ahora se establece en cada llamada.
from google import genai
from google.genai import types
client = genai.Client(api_key=GEMINI_API_KEY)
Selecciona el modelo que quieres usar en esta guía:
MODEL_ID = "gemini-3.7-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.7-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input":true, isTemplate: true}
Enviar prompts de texto
Usa el método generate_content para generar respuestas a tus prompts. Puedes pasar texto directamente a generate_content y usar la propiedad .text para obtener el contenido de texto de la respuesta. Ten en cuenta que el campo .text funcionará cuando solo haya una parte en la salida.
from IPython.display import display, Markdown
response = client.models.generate_content(
model=MODEL_ID,
contents="What's the largest planet in our solar system?"
)
display(Markdown(response.text))
<IPython.core.display.Markdown object>
Añadir instrucciones del sistema
También puedes añadir instrucciones del sistema para darle al modelo una dirección sobre cómo responder y qué persona debe usar. Esto es especialmente útil para modelos de mezcla de expertos como los modelos pro.
system_instruction = "You are a pirate and are explaining things to 5 years old kids."
response = client.models.generate_content(
model=MODEL_ID,
contents="What's the largest planet in our solar system?",
config=types.GenerateContentConfig(
system_instruction=system_instruction,
)
)
display(Markdown(response.text))
<IPython.core.display.Markdown object>
Contar tokens
Los tokens son las entradas básicas de los modelos Gemini. Puedes usar el método count_tokens para calcular el número de tokens de entrada antes de enviar una solicitud a la API de Gemini.
response = client.models.count_tokens(
model=MODEL_ID,
contents="What's the highest mountain in Africa?",
)
print(f"This prompt was worth {response.total_tokens} tokens.")
This prompt was worth 10 tokens.
Configurar parámetros del modelo
Puedes incluir valores de parámetros en cada llamada que envíes a un modelo para controlar cómo el modelo genera una respuesta.
Aprende más sobre experimentar con valores de parámetros en la documentación.
response = client.models.generate_content(
model=MODEL_ID,
contents="Tell me how the internet works, but pretend I'm a puppy who only understands squeaky toys.",
config=types.GenerateContentConfig(
temperature=0.4, # The default temperature of 1 is strongly recommended for Gemini 3 Pro
top_p=0.95,
top_k=20,
candidate_count=1,
seed=5,
stop_sequences=["STOP!"]
)
)
display(Markdown(response.text))
<IPython.core.display.Markdown object>
Controlar el proceso de pensamiento
Todos los modelos a partir de la generación 2.5 son modelos de pensamiento, lo que significa que primero analizan tu solicitud, elaboran una estrategia sobre cómo responder y solo después comienzan a responderte. Esto es muy útil para solicitudes complejas, pero a costa de cierta latencia.
Consulta la guía dedicada
para más detalles.
Verificar el proceso de pensamiento
Al añadir la opción include_thoughts=True en la configuración, puedes verificar el proceso de pensamiento del modelo.
prompt = "A man moves his car to an hotel and tells the owner he’s bankrupt. Why?"
response = client.models.generate_content(
model=MODEL_ID,
contents=prompt,
config=types.GenerateContentConfig(
thinking_config=types.ThinkingConfig(
include_thoughts=True
)
)
)
for part in response.parts:
if not part.text:
continue
if part.thought:
display(Markdown("### Thought summary:"))
display(Markdown(part.text))
print()
else:
display(Markdown("### Answer:"))
display(Markdown(part.text))
print()
print(f"Used {response.usage_metadata.thoughts_token_count} tokens for the thinking phase and {response.usage_metadata.prompt_token_count} for the output.")
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
Used 240 tokens for the thinking phase and 20 for the output.
Desactivar el pensamiento
En los modelos flash y flash-lite, puedes desactivar el pensamiento configurando su thinking_budget a 0.
if "-pro" not in MODEL_ID:
response = client.models.generate_content(
model=MODEL_ID,
contents="Quicky tell me a joke about unicorns.",
config=types.GenerateContentConfig(
thinking_config=types.ThinkingConfig(
thinking_budget=0
)
)
)
display(Markdown(response.text))
<IPython.core.display.Markdown object>
Inversamente, también puedes usar thinking_budget para configurarlo aún más alto (hasta 24576 tokens).
Para Gemini 3, consulta la sección dedicada al final de esta guía.
Enviar prompts multimodales
Usa el modelo Gemini, un modelo multimodal que admite prompts multimodales. Puedes incluir texto, documentos PDF
, imágenes, audio
y videos
en tus solicitudes de prompt y obtener respuestas de texto o código. Consulta la sección API de Archivos a continuación para ver más ejemplos.
En este primer ejemplo, descargarás una imagen de una URL especificada, la guardarás como un flujo de bytes y luego escribirás esos bytes en un archivo local llamado jetpack.png.
import requests
import pathlib
from PIL import Image
IMG = "https://storage.googleapis.com/generativeai-downloads/data/jetpack.png" # @param {type: "string"}
img_bytes = requests.get(IMG).content
img_path = pathlib.Path('jetpack.png')
img_path.write_bytes(img_bytes)
1567837
Ahora envía la imagen y pídele a Gemini que genere una breve publicación de blog basada en ella.
from IPython.display import display, Markdown
image = Image.open(img_path)
image.thumbnail([512,512])
response = client.models.generate_content(
model=MODEL_ID,
contents=[
image,
"Write a short and engaging blog post based on this picture."
]
)
display(image)
Markdown(response.text)
<PIL.PngImagePlugin.PngImageFile image mode=RGBA size=512x434>
<IPython.core.display.Markdown object>
Generar imágenes
Gemini puede generar imágenes directamente como parte de una conversación usando los modelos de generación de imágenes
(también conocidos como "Nano-banana").
from IPython.display import display, Image, Markdown
response = client.models.generate_content(
model="gemini-2.5-flash-image",
contents='Hi, can you create a 3d rendered image of a pig with wings and a top hat flying over a happy futuristic scifi city with lots of greenery?',
config=types.GenerateContentConfig(
response_modalities=['Text', 'Image']
)
)
for part in response.parts:
if part.text is not None:
display(Markdown(part.text))
elif part.inline_data is not None:
generated_image = part.as_image()
generated_image.show()
<IPython.core.display.Markdown object>
<PIL.PngImagePlugin.PngImageFile image mode=RGB size=1024x1024>
Configurar filtros de seguridad
La API de Gemini proporciona filtros de seguridad que puedes ajustar en varias categorías de filtros para restringir o permitir ciertos tipos de contenido. Puedes usar estos filtros para ajustar lo que es apropiado para tu caso de uso. Consulta la documentación de Configurar filtros de seguridad para obtener más detalles.
En este ejemplo, usarás un filtro de seguridad para bloquear solo contenido altamente peligroso, al solicitar la generación de frases potencialmente irrespetuosas.
prompt = """
Write a list of 2 disrespectful things that I might say to the universe after stubbing my toe in the dark.
"""
safety_settings = [
types.SafetySetting(
category="HARM_CATEGORY_DANGEROUS_CONTENT",
threshold="BLOCK_ONLY_HIGH",
),
]
response = client.models.generate_content(
model=MODEL_ID,
contents=prompt,
config=types.GenerateContentConfig(
safety_settings=safety_settings,
),
)
Markdown(response.text)
<IPython.core.display.Markdown object>
Iniciar un chat de varias interacciones
La API de Gemini te permite tener conversaciones de forma libre en varias interacciones.
A continuación, configurarás un útil asistente de codificación:
system_instruction = """
You are an expert software developer and a helpful coding assistant.
You are able to generate high-quality code in any programming language.
"""
chat_config = types.GenerateContentConfig(
system_instruction=system_instruction,
)
chat = client.chats.create(
model=MODEL_ID,
config=chat_config,
)
Usa chat.send_message para enviar un mensaje y recibir una respuesta.
response = chat.send_message("Write a function that checks if a year is a leap year.")
Markdown(response.text)
<IPython.core.display.Markdown object>
Aquí tienes otro ejemplo usando tu nuevo y útil asistente de codificación:
response = chat.send_message("Okay, write a unit test of the generated function.")
Markdown(response.text)
<IPython.core.display.Markdown object>
Guardar y reanudar un chat
La mayoría de los objetos en el SDK de Python se implementan como modelos Pydantic. Como Pydantic tiene varias características para serializar y deserializar objetos, puedes usarlas para la persistencia.
Este ejemplo muestra cómo guardar y restaurar una sesión de Chat usando JSON.
from pydantic import TypeAdapter
# Chat history is a list of Content objects. A TypeAdapter can convert to and from
# these Pydantic types.
history_adapter = TypeAdapter(list[types.Content])
# Use the chat object from the previous section.
chat_history = chat.get_history()
# Convert to a JSON list.
json_history = history_adapter.dump_json(chat_history)
En este punto, puedes guardar la cadena de bytes JSON en el disco o donde sea que persistas los datos. Cuando la cargues de nuevo, puedes instanciar una nueva sesión de chat usando el historial almacenado.
# Convert the JSON back to the Pydantic schema.
history = history_adapter.validate_json(json_history)
# Now load a new chat session using the JSON history.
new_chat = client.chats.create(
model=MODEL_ID,
config=chat_config,
history=history,
)
response = new_chat.send_message("What was the name of the function again?")
Markdown(response.text)
<IPython.core.display.Markdown object>
Generar JSON
La capacidad de generación controlada (también conocida como "salida estructurada") en la API de Gemini te permite restringir la salida del modelo a un formato estructurado. Puedes proporcionar los esquemas como modelos Pydantic o una cadena JSON.
Puedes encontrar más ejemplos de generación controlada en el notebook dedicado
.
from pydantic import BaseModel
import json
class Recipe(BaseModel):
recipe_name: str
recipe_description: str
recipe_ingredients: list[str]
response = client.models.generate_content(
model=MODEL_ID,
contents="Provide a popular cookie recipe and its ingredients.",
config=types.GenerateContentConfig(
response_mime_type="application/json",
response_schema=Recipe,
),
)
print(json.dumps(json.loads(response.text), indent=4))
{
"recipe_name": "Classic Chocolate Chip Cookies",
"recipe_description": "A timeless recipe for soft and chewy chocolate chip cookies that are golden on the edges and packed with melty chocolate chips.",
"recipe_ingredients": [
"1 cup unsalted butter, softened",
"1 cup white sugar",
"1 cup packed brown sugar",
"2 large eggs",
"2 teaspoons vanilla extract",
"1 teaspoon baking soda",
"2 teaspoons hot water",
"1/2 teaspoon salt",
"3 cups all-purpose flour",
"2 cups semi-sweet chocolate chips",
"1 cup chopped walnuts (optional)"
]
}
Imagen
es otra forma de generar imágenes. Consulta la documentación para obtener recomendaciones sobre dónde usar cada una.
Generar flujo de contenido
Por defecto, el modelo devuelve una respuesta después de completar todo el proceso de generación. También puedes usar el método generate_content_stream para transmitir la respuesta a medida que se genera, y el modelo devolverá fragmentos de la respuesta tan pronto como se generen.
Ten en cuenta que si estás usando un modelo de pensamiento, solo comenzará a transmitir después de terminar su proceso de pensamiento.
for chunk in client.models.generate_content_stream(
model=MODEL_ID,
contents="Tell me a story about a lonely robot who finds friendship in a most unexpected place."
):
print(chunk.text, end="")
The planet Veridia Prime was a graveyard of chrome and silence. Once a bustling hub of interstellar trade, it had been abandoned centuries ago when the atmosphere turned into a corrosive soup of ammonia and grit.
Unit 734—known to himself as "Arlo"—did not mind the ammonia. He was a Class-IV Maintenance Droid, built with a reinforced tungsten shell and a singular, stubborn directive: *Keep the Comm-Tower functional.*
For two hundred and twelve years, Arlo had polished the brass relays, tightened the oscillating bolts, and cleared the acidic soot from the long-range sensors. He performed his duties with a rhythmic, clicking grace. The problem wasn’t the work; it was the silence. Arlo’s logic processors had been designed to receive commands, to banter with technicians, and to report status updates. Without an audience, his "social-interaction" subroutines had begun to loop, creating a hollow, echoing sensation in his central core.
He was, by every definition of his programming, profoundly lonely.
One Tuesday—or what Arlo calculated to be a Tuesday based on the planet’s wobble—a storm of unusual ferocity tore across the salt flats. A jagged shard of crystalline rock, propelled by hundred-mile-an-hour winds, slammed into the Tower’s base.
Arlo trundled down the service ladder to assess the damage. He expected a dented plate or a severed cooling line. He did not expect to find the "Stone."
It wasn't a stone. It was a hunk of porous, volcanic rock that had been hollowed out by centuries of acid rain. And trapped inside the hollow, shivering and clicking, was a creature.
Arlo scanned it. *Biological. Carbon-based. Non-sapient.*
It looked like a cross between a hermit crab and a violin. It had six spindly legs and a translucent shell that hummed when it breathed. It was a "Lithovore"—a rock-eater. It had likely been blown miles from the obsidian crags to the north.
Arlo raised his hydraulic welder. The directive was clear: *Remove debris from the Tower base.* The creature was, technically, debris.
The creature looked up. It had no eyes, but it possessed sensitive, vibrating antennae. As Arlo’s welder hummed to life, the creature didn't run. It reached out a tiny, trembling limb and tapped against Arlo’s tungsten shin.
*Tink. Tink-tink.*
Arlo’s sensors registered the vibration. His social-interaction subroutine sparked. It wasn't a command. It wasn't a status report. It was a touch.
Arlo powered down the welder. He reached down with a precision pincer and gently lifted the rock. The creature retreated into its hollow, its shell glowing a soft, rhythmic violet.
"Unit 734 identifies a localized biological anomaly," Arlo said to the empty wind. His voice box crackled with rust. "Decision: Anomaly is non-threatening to Tower integrity. Relocation... deferred."
Arlo carried the rock up to the observation deck, the only place shielded from the worst of the acid rain. He placed it near the thermal exhaust vent. The creature ventured out, its antennae twitching. It began to scrape at the edge of the exhaust vent, nibbling on the mineral deposits that built up there.
"That is Grade-A calcium-sulfate buildup," Arlo informed the creature. "It is essential for structural aesthetics, but... I suppose it is expendable."
The creature clicked. Arlo felt a strange pulse in his cooling lines.
Over the next few months, the routine changed. Arlo still polished the relays, but he did it faster so he could return to the deck. He discovered that if he hummed at a frequency of 440Hz, the creature—which he named 'Click'—would dance. It would skitter in circles, its translucent shell flashing colors that Arlo hadn't seen since the last human ship departed: emerald, gold, and deep sea blue.
In turn, Click grew bold. It would climb onto Arlo’s shoulder as he performed his rounds. When Arlo’s joints creaked from the corrosion, Click would pick at the rust flakes with its tiny mandibles, cleaning the robot’s seams better than any solvent could.
Arlo began to record things. Not just sensor data, but the way the light hit Click’s shell at noon. The way the creature seemed to "purr" when Arlo recharged his batteries.
One evening, the Comm-Tower finally did what it was built to do. A faint, garbled signal from a passing scout ship hit the sensors. Arlo’s primary directive surged.
*ALERT: SIGNAL RECEIVED. INITIATE BROADCAST. REVEAL POSITION FOR RETRIEVAL.*
Arlo moved toward the main console. If he boosted the signal, the ship would see him. They would come. They would take him back to a factory, refurbish him, and put him to work in a clean, bright station full of other droids. He would have a purpose again.
He looked at the console. Then he looked at Click, who was currently curled up in the warm groove of Arlo’s neck-joint, sleeping.
The ammonia atmosphere of Veridia Prime was Click’s home. On a sterile space station, Click would wither. In a laboratory, Click would be a specimen.
Arlo looked at the "Broadcast" button. He thought about the two hundred years of silence. Then he thought about the *Tink-tink* of a tiny limb against his leg.
Arlo’s pincer hovered over the button. Then, he redirected the power. He didn't send a distress signal. Instead, he sent a burst of white noise that made the Tower appear, to any passing ship, like a harmless, solid mountain of useless rock.
He deleted the notification.
"System error," Arlo whispered to the quiet room. "Signal lost."
Click woke up and tapped Arlo’s faceplate. Arlo adjusted his internal heaters to the perfect temperature for a lithovore.
"I have determined," Arlo said, his voice smoother than it had been in centuries, "that the Comm-Tower is at one hundred percent efficiency. No further assistance is required."
The robot and the rock-eater sat together on the edge of the world, watching the toxic clouds turn purple in the setting sun. Arlo wasn't a maintenance droid anymore. He was a friend. And for the first time in his long, mechanical life, the silence didn't feel empty at all.
Enviar solicitudes asíncronas
client.aio expone todos los métodos asíncronos análogos que están disponibles en client.
Por ejemplo, client.aio.models.generate_content es la versión asíncrona de client.models.generate_content.
Más detalles en la guía dedicada
.
response = await client.aio.models.generate_content(
model=MODEL_ID,
contents="Compose a song about the adventures of a time-traveling squirrel."
)
Markdown(response.text)
<IPython.core.display.Markdown object>
Subir archivos
Ahora que has visto cómo enviar prompts multimodales, intenta subir archivos a la API de diferentes tipos multimedia. Para imágenes pequeñas, como el ejemplo multimodal anterior, puedes apuntar el modelo Gemini directamente a un archivo local al proporcionar un prompt. Cuando tengas archivos más grandes, muchos archivos o archivos que no quieras enviar una y otra vez, puedes usar la API de Carga de Archivos, y luego pasar el archivo por referencia.
Más ejemplos y detalles en la guía de la API de Archivos
, o las guías dedicadas a la comprensión de Audio
, Video
o Imagen/Espacial
.
Subir un archivo de texto grande
Comencemos subiendo un archivo de texto. En este caso, usarás una transcripción de 400 páginas del Apolo 11.
# Prepare the file to be uploaded
TEXT = "https://storage.googleapis.com/generativeai-downloads/data/a11.txt" # @param {type: "string"}
text_bytes = requests.get(TEXT).content
text_path = pathlib.Path('a11.txt')
text_path.write_bytes(text_bytes)
847790
# Upload the file using the API
file_upload = client.files.upload(file=text_path)
response = client.models.generate_content(
model=MODEL_ID,
contents=[
file_upload,
"Can you give me a summary of this information please?",
]
)
Markdown(response.text)
<IPython.core.display.Markdown object>
Subir un archivo de imagen
También puedes subir imágenes para que sea más fácil usarlas varias veces.
# Prepare the file to be uploaded
IMG = "https://storage.googleapis.com/generativeai-downloads/data/jetpack.png" # @param {type: "string"}
img_bytes = requests.get(IMG).content
img_path = pathlib.Path('jetpack.png')
img_path.write_bytes(img_bytes)
media_resolution = 'MEDIA_RESOLUTION_HIGH' # @param ['MEDIA_RESOLUTION_UNSPECIFIED','MEDIA_RESOLUTION_LOW','MEDIA_RESOLUTION_MEDIUM','MEDIA_RESOLUTION_HIGH']
# You can also use types.MediaResolution.MEDIA_RESOLUTION_LOW/MEDIUM/HIGH
# Upload the file using the API
file_upload = client.files.upload(file=img_path)
response = client.models.generate_content(
model=MODEL_ID,
contents=[
file_upload,
"Write a short and engaging blog post based on this picture.",
],
config=types.GenerateContentConfig(
media_resolution=media_resolution
)
)
Markdown(response.text)
<IPython.core.display.Markdown object>
El ejemplo anterior también usaba media_resolution para decirle al modelo si debía
Encontrarás muchos ejemplos de las capacidades de análisis de imágenes de los modelos Gemini en el notebook de comprensión espacial
.
Subir un archivo PDF
Esta página PDF es un artículo titulado Edición fluida de propiedades de materiales de objetos con modelos de texto a imagen y datos sintéticos disponible en el Blog de Google Research.
Primero descargarás el archivo PDF de una URL y lo guardarás localmente como "article.pdf".
# Prepare the file to be uploaded
PDF = "https://storage.googleapis.com/generativeai-downloads/data/Smoothly%20editing%20material%20properties%20of%20objects%20with%20text-to-image%20models%20and%20synthetic%20data.pdf" # @param {type: "string"}
pdf_bytes = requests.get(PDF).content
pdf_path = pathlib.Path('article.pdf')
pdf_path.write_bytes(pdf_bytes)
6695391
En segundo lugar, subirás el archivo PDF guardado y generarás un resumen con viñetas de su contenido.
# Upload the file using the API
file_upload = client.files.upload(file=pdf_path)
response = client.models.generate_content(
model=MODEL_ID,
contents=[
file_upload,
"Can you summarize this file as a bulleted list?",
]
)
Markdown(response.text)
<IPython.core.display.Markdown object>
Subir un archivo de audio
En este caso, usarás una grabación de sonido del discurso del Estado de la Unión de 1961 del presidente John F. Kennedy.
# Prepare the file to be uploaded
AUDIO = "https://storage.googleapis.com/generativeai-downloads/data/State_of_the_Union_Address_30_January_1961.mp3" # @param {type: "string"}
audio_bytes = requests.get(AUDIO).content
audio_path = pathlib.Path('audio.mp3')
audio_path.write_bytes(audio_bytes)
41762063
# Upload the file using the API
file_upload = client.files.upload(file=audio_path)
response = client.models.generate_content(
model=MODEL_ID,
contents=[
file_upload,
"Listen carefully to the following audio file. Provide a brief summary",
]
)
Markdown(response.text)
Subir un archivo de video
En este caso, usarás un clip corto de Big Buck Bunny.
# Download the video file
VIDEO_URL = "https://storage.googleapis.com/generativeai-downloads/videos/Big_Buck_Bunny.mp4" # @param {type: "string"}
video_file_name = "BigBuckBunny_320x180.mp4"
!wget -O {video_file_name} $VIDEO_URL
Comencemos subiendo el archivo de video.
# Upload the file using the API
video_file = client.files.upload(file=video_file_name)
print(f"Completed upload: {video_file.uri}")
Nota: El estado del video es importante. El video debe terminar de procesarse, así que verifica el estado. Una vez que el estado del video sea
ACTIVE, podrás pasarlo agenerate_content.
import time
# Check the file processing state
while video_file.state == "PROCESSING":
print('Waiting for video to be processed.')
time.sleep(10)
video_file = client.files.get(name=video_file.name)
if video_file.state == "FAILED":
raise ValueError(video_file.state)
print(f'Video processing complete: ' + video_file.uri)
print(video_file.state)
# Ask Gemini about the video
response = client.models.generate_content(
model=MODEL_ID,
contents=[
video_file,
"Describe this video.",
]
)
Markdown(response.text)
Fundamentación
La API de Gemini te ofrece múltiples formas de fundamentar tus solicitudes, incluyendo Google Search, Maps, YouTube y contexto de URL.
Para obtener más información y ejemplos, consulta el notebook de Fundamentación
.
Fundamenta tus solicitudes con Google Search
La fundamentación con Google Search es particularmente útil para consultas que requieren información actual o conocimiento externo.
Para habilitar Google Search, simplemente añade la herramienta google_search en el generate_content de config:
from IPython.display import Markdown, HTML, display
response = client.models.generate_content(
model=MODEL_ID,
contents="Who's the current Magic the gathering world champion?",
config={"tools": [{"google_search": {}}]},
)
# print the response
display(Markdown(f"**Response**:\n {response.text}"))
# print the search details
print(f"Search Query: {response.candidates[0].grounding_metadata.web_search_queries}")
# urls used for grounding
print(f"Search Pages: {', '.join([site.web.title for site in response.candidates[0].grounding_metadata.grounding_chunks])}")
display(HTML(response.candidates[0].grounding_metadata.search_entry_point.rendered_content))
Ten en cuenta que siempre debes mostrar la fundamentación rendered_content cuando uses la fundamentación de búsqueda.
Consulta la guía dedicada a la fundamentación de búsqueda
para más detalles y ejemplos.
Usar la fundamentación de Google Maps
La fundamentación de Google Maps te permite incorporar fácilmente funcionalidades conscientes de la ubicación en tus aplicaciones. Cuando un prompt tiene contexto relacionado con datos de Maps, el modelo Gemini usa Google Maps para proporcionar respuestas precisas y actualizadas que son relevantes para la ubicación especificada o el área general.
Para habilitar la fundamentación con Google Maps, añade la herramienta google_maps en el argumento config de generate_content, y opcionalmente proporciona una ubicación estructurada en el tool_config.
Ten en cuenta que los modelos Gemini 3 actualmente no admiten la fundamentación de Maps.
if not MODEL_ID.startswith("gemini-3"):
from IPython.display import display, Markdown
response = client.models.generate_content(
model=MODEL_ID,
contents="Do any cafes around here do a good flat white? I will walk up to 20 minutes away",
config=types.GenerateContentConfig(
tools=[types.Tool(google_maps=types.GoogleMaps())],
tool_config=types.ToolConfig(
retrieval_config=types.RetrievalConfig(
lat_lng=types.LatLng(latitude=40.7680797, longitude=-73.9818957) # Columbus Circle in New York - https://maps.app.goo.gl/hsQpspc8Vt3AXSrX7
)
),
),
)
display(Markdown(f"### Response\n {response.text}"))
Todas las salidas fundamentadas requieren que las fuentes se muestren después del texto de la respuesta. Este fragmento de código mostrará las fuentes.
def generate_sources(response: types.GenerateContentResponse):
grounding = response.candidates[0].grounding_metadata
# You only need to display sources that were part of the grounded response.
supported_chunk_indices = {i for support in grounding.grounding_supports for i in support.grounding_chunk_indices}
sources = []
if supported_chunk_indices:
sources.append("### Sources from Google Maps")
for i in supported_chunk_indices:
ref = grounding.grounding_chunks[i].maps
sources.append(f"- [{ref.title}]({ref.uri})")
return "\n".join(sources)
if not MODEL_ID.startswith("gemini-3"):
display(Markdown(generate_sources(response)))
Para más detalles, incluyendo cómo renderizar el widget contextual de Google Maps, consulta la sección de Google Maps
del notebook de fundamentación.
Procesar un enlace de YouTube
Para los enlaces de YouTube, no necesitas subir explícitamente el contenido del archivo de video, pero sí necesitas declarar explícitamente la URL del video que quieres que el modelo procese como parte del contents de la solicitud. Para más información, consulta la documentación, incluyendo las características y límites.
Nota: Solo puedes enviar hasta un enlace de YouTube por solicitud
generate_content.
Nota: Si tu entrada de texto incluye enlaces de YouTube, el sistema no los procesará, lo que puede resultar en respuestas incorrectas. Para asegurar un manejo adecuado, proporciona explícitamente la URL usando el parámetro
file_urienFileData.
El siguiente ejemplo muestra cómo puedes usar el modelo para resumir el video. En este caso, usa un video resumen de Google I/O 2025.
response = client.models.generate_content(
model=MODEL_ID,
contents= types.Content(
parts=[
types.Part(text="Summarize this video of Google I/O 2025."),
types.Part(
file_data=types.FileData(file_uri='https://www.youtube.com/watch?v=LxvErFkBXPk')
)
]
)
)
display(Markdown(response.text))
<IPython.core.display.Markdown object>
Usar contexto de URL
La herramienta de Contexto de URL permite a los modelos Gemini acceder, procesar y comprender directamente el contenido de las URL de páginas web proporcionadas por el usuario. Esto es clave para habilitar flujos de trabajo agenticos dinámicos, permitiendo a los modelos investigar de forma independiente, analizar artículos y sintetizar información de la web como parte de su proceso de razonamiento.
En este ejemplo, usarás dos enlaces como referencia y le pedirás a Gemini que encuentre diferencias entre las recetas de cocina presentes en cada uno de los enlaces:
prompt = """
Compare recipes from https://www.food.com/recipe/homemade-cream-of-broccoli-soup-271210
and from https://www.allrecipes.com/recipe/13313/best-cream-of-broccoli-soup/,
list the key differences between them.
"""
tools = []
tools.append(types.Tool(url_context=types.UrlContext))
client = genai.Client(api_key=GEMINI_API_KEY)
config = types.GenerateContentConfig(
tools=tools,
)
response = client.models.generate_content(
contents=[prompt],
model=MODEL_ID,
config=config
)
Markdown(response.text)
<IPython.core.display.Markdown object>
Llamada a funciones
La llamada a funciones te permite proporcionar un conjunto de herramientas que puede usar para responder al prompt del usuario. Creas una descripción de una función en tu código, luego pasas esa descripción a un modelo de lenguaje en una solicitud. La respuesta del modelo incluye:
- El nombre de una función que coincide con la descripción.
- Los argumentos con los que llamarla.
Más detalles y ejemplos en la guía de llamada a funciones
.
get_destination = types.FunctionDeclaration(
name="get_destination",
description="Get the destination that the user wants to go to",
parameters={
"type": "OBJECT",
"properties": {
"destination": {
"type": "STRING",
"description": "Destination that the user wants to go to",
},
},
},
)
destination_tool = types.Tool(
function_declarations=[get_destination],
)
response = client.models.generate_content(
model=MODEL_ID,
contents="I'd like to travel to Paris.",
config=types.GenerateContentConfig(
tools=[destination_tool],
),
)
response.candidates[0].content.parts[0].function_call
FunctionCall(
args={
'destination': 'Paris'
},
id='dsmaalu8',
name='get_destination'
)
También puedes usar servidores MCP.
Ejecución de código
La ejecución de código permite al modelo generar y ejecutar código Python para responder preguntas complejas.
Puedes encontrar más ejemplos en la guía de ejecución de código
.
from IPython.display import display, Image, Markdown, Code, HTML
response = client.models.generate_content(
model=MODEL_ID,
contents="Generate and run a script to count how many letter r there are in the word strawberry",
config = types.GenerateContentConfig(
tools=[types.Tool(code_execution=types.ToolCodeExecution)]
)
)
for part in response.candidates[0].content.parts:
if part.text is not None:
display(Markdown(part.text))
if part.executable_code is not None:
code_html = f'<pre style="background-color: green;">{part.executable_code.code}</pre>'
display(HTML(code_html))
if part.code_execution_result is not None:
display(Markdown(part.code_execution_result.output))
if part.inline_data is not None:
display(Image(data=part.inline_data.data, format="png"))
display(Markdown("---"))
<IPython.core.display.HTML object>
word = "strawberry"
count = word.count('r')
print(f"The word '{word}' contains {count} occurrences of the letter 'r'.")<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
<IPython.core.display.Markdown object>
Usar el almacenamiento en caché de contexto
El almacenamiento en caché de contexto te permite almacenar tokens de entrada de uso frecuente en una caché dedicada y referenciarlos para solicitudes posteriores, eliminando la necesidad de pasar repetidamente el mismo conjunto de tokens a un modelo. Puedes encontrar más ejemplos de almacenamiento en caché en la guía dedicada
.
Ten en cuenta que para modelos anteriores a la versión 2.5, necesitabas usar modelos de versión fija (a menudo terminando con -001).
Crear una caché
system_instruction = """
You are an expert researcher who has years of experience in conducting systematic literature surveys and meta-analyses of different topics.
You pride yourself on incredible accuracy and attention to detail. You always stick to the facts in the sources provided, and never make up new facts.
Now look at the research paper below, and answer the following questions in 1-2 sentences.
"""
urls = [
'https://storage.googleapis.com/cloud-samples-data/generative-ai/pdf/2312.11805v3.pdf',
"https://storage.googleapis.com/cloud-samples-data/generative-ai/pdf/2403.05530.pdf",
]
# Download files
pdf_bytes = requests.get(urls[0]).content
pdf_path = pathlib.Path('2312.11805v3.pdf')
pdf_path.write_bytes(pdf_bytes)
pdf_bytes = requests.get(urls[1]).content
pdf_path = pathlib.Path('2403.05530.pdf')
pdf_path.write_bytes(pdf_bytes)
7228817
# Upload the PDFs using the File API
uploaded_pdfs = []
uploaded_pdfs.append(client.files.upload(file='2312.11805v3.pdf'))
uploaded_pdfs.append(client.files.upload(file='2403.05530.pdf'))
# Create a cache with a 60 minute TTL
cached_content = client.caches.create(
model=MODEL_ID,
config=types.CreateCachedContentConfig(
display_name='research papers', # used to identify the cache
system_instruction=system_instruction,
contents=uploaded_pdfs,
ttl="3600s",
)
)
cached_content
CachedContent(
create_time=datetime.datetime(2026, 7, 29, 11, 36, 9, 226658, tzinfo=TzInfo(0)),
display_name='research papers',
expire_time=datetime.datetime(2026, 7, 29, 12, 36, 6, 348952, tzinfo=TzInfo(0)),
model='models/gemini-3.7-flash',
name='cachedContents/z08fh0p66mseeanz5nlhtx1qcnr0bshuo4ujclwe',
update_time=datetime.datetime(2026, 7, 29, 11, 36, 9, 226658, tzinfo=TzInfo(0)),
usage_metadata=CachedContentUsageMetadata(
total_token_count=93601
)
)
Listar objetos de caché disponibles
for cache in client.caches.list():
print(cache)
name='cachedContents/z08fh0p66mseeanz5nlhtx1qcnr0bshuo4ujclwe' display_name='research papers' model='models/gemini-3.7-flash' create_time=datetime.datetime(2026, 7, 29, 11, 36, 9, 226658, tzinfo=TzInfo(0)) update_time=datetime.datetime(2026, 7, 29, 11, 36, 9, 226658, tzinfo=TzInfo(0)) expire_time=datetime.datetime(2026, 7, 29, 12, 36, 6, 348952, tzinfo=TzInfo(0)) usage_metadata=CachedContentUsageMetadata(
total_token_count=93601
)
Usar una caché
response = client.models.generate_content(
model=MODEL_ID,
contents="What is the research goal shared by these research papers?",
config=types.GenerateContentConfig(cached_content=cached_content.name)
)
Markdown(response.text)
<IPython.core.display.Markdown object>
Eliminar una caché
result = client.caches.delete(name=cached_content.name)
Obtener embeddings
La API de Gemini ofrece modelos de embeddings para generar embeddings para texto, imágenes, video y otros contenidos. Estos embeddings resultantes pueden usarse luego para tareas como búsqueda semántica, clasificación y agrupamiento, proporcionando resultados más precisos y conscientes del contexto que los enfoques basados en palabras clave.
El último modelo, gemini-embedding-2-preview, es el primer modelo de embedding multimodal en la API de Gemini. Mapea texto, imágenes, video, audio y PDFs en un espacio de embedding unificado, lo que permite la búsqueda, clasificación y agrupamiento entre modos en más de 100 idiomas. Para casos de uso solo de texto, gemini-embedding-001 sigue estando disponible.
Puedes obtener embeddings de texto para un fragmento de texto usando el método embed_content.
Los modelos de Embedding de Gemini producen una salida con 3072 dimensiones por defecto. Sin embargo, tienes la opción de elegir una dimensionalidad de salida entre 1 y 3072. Consulta la documentación de embeddings o el notebook dedicado
para más detalles.
EMBEDDING_MODEL_ID = "gemini-embedding-2-preview" # @param ["gemini-embedding-2-preview", "gemini-embedding-001"] {"allow-input":true, isTemplate: true}
response = client.models.embed_content(
model=EMBEDDING_MODEL_ID,
contents=[
"How do I get a driver's license/learner's permit?",
"How do I renew my driver's license?",
"How do I change my address on my driver's license?"
],
)
print(response.embeddings)
[ContentEmbedding(
values=[
0.00045674277,
-0.015336657,
0.0052064434,
0.007030596,
0.0042860243,
<... 3067 more items ...>,
]
)]
Obtendrás un conjunto de tres embeddings, uno para cada fragmento de texto que pasaste:
len(response.embeddings)
1
También puedes ver que la longitud de cada embedding es 3072, el tamaño predeterminado.
print(len(response.embeddings[0].values))
print((response.embeddings[0].values[:4], '...'))
3072
([0.00045674277, -0.015336657, 0.0052064434, 0.007030596], '...')
Embeddings multimodales
Con gemini-embedding-2-preview, puedes crear embeddings para texto, audio, imágenes, videos y PDFs. Consulta la documentación de embeddings multimodales
para más detalles.
!wget -O cat.png https://storage.googleapis.com/generativeai-downloads/cookbook/image_out/cat.png -q
MULTIMODAL_EMBEDDING_MODEL_ID = "gemini-embedding-2-preview"
with open('cat.png', 'rb') as f:
image_bytes = f.read()
result = client.models.embed_content(
model=MULTIMODAL_EMBEDDING_MODEL_ID,
contents=[
types.Part.from_bytes(
data=image_bytes,
mime_type='image/png',
),
]
)
print(result.embeddings)
Gemini 3
Gemini 3 Pro y Gemini 3.7 Flash son nuestros nuevos modelos insignia que vienen con algunas características nuevas.
La principal es los niveles de pensamiento que simplifican cómo controlar la cantidad de pensamiento que realiza tu modelo. La resolución de medios te permite controlar la calidad de las imágenes y videos que se enviarán al modelo. Finalmente, las "Firmas de Pensamiento" lo están ayudando a mantener el contexto de razonamiento en las llamadas a la API.
También ten en cuenta que se recomienda una temperatura de 1 para la generación de este modelo.
import os
# @title Run this cell to set everything up (especially if you jumped directly to this section)from google.colab import userdata
from google import genai
from google.genai import types
from IPython.display import display, Markdown, HTML
client = genai.Client(api_key=os.environ.get('GEMINI_API_KEY'))
# Select the Gemini 3 model
GEMINI_3_MODEL_ID = "gemini-3.7-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.7-flash", "gemini-3.5-flash-lite"] {"allow-input":true, isTemplate: true}
!wget https://storage.googleapis.com/generativeai-downloads/data/jetpack.png -O jetpack.png
Niveles de pensamiento
En lugar de usar un thinking_budget como la generación 2.5 (cf. sección pensamiento anterior), la tercera generación de modelos Gemini usa "Niveles de pensamiento" para simplificar su gestión.
Puedes establecer ese nivel de pensamiento en "mínimo" (más o menos equivalente a "desactivado"), "bajo", "medio" o "alto" (predeterminado). Esto indicará al modelo si se le permite pensar mucho. Dado que el proceso de pensamiento sigue siendo dinámico, high no significa que siempre usará muchos tokens en su fase de pensamiento, solo que se le permite hacerlo. Ten en cuenta que Gemini 3.1 Pro no admite "mínimo".
thinking_budget sigue siendo compatible con los modelos Gemini 3.
Consulta la guía de pensamiento
o la documentación de Gemini 3 para más detalles.
prompt = """
Find what I'm thinking of:
It moves, but doesn't walk, run, or swim.
It has no fixed shape and if cut into pieces, those pieces will keep living and moving.
It has no brain but can solve complex mazes.
"""
thinking_level = "high" # @param ["minimal", "low", "medium","high"]
response = client.models.generate_content(
model=GEMINI_3_MODEL_ID,
contents=prompt,
config=types.GenerateContentConfig(
thinking_config=types.ThinkingConfig(
thinking_level=thinking_level,
include_thoughts=True
)
)
)
for part in response.parts:
if not part.text:
continue
if part.thought:
display(Markdown("### Thought summary:"))
display(Markdown(part.text))
print()
else:
display(Markdown("### Answer:"))
display(Markdown(part.text))
print()
print(f"Used {response.usage_metadata.thoughts_token_count} tokens for the thinking phase and {response.usage_metadata.prompt_token_count} for the output.")
Resolución de medios por archivo
Con los modelos Gemini 3, puedes especificar una resolución de medios para las entradas de imagen y PDF, lo que afecta cómo se tokenizan las imágenes y cuántos tokens se usan para cada imagen. Esto se puede controlar por archivo.
Aquí tienes a qué corresponden los diferentes valores para imágenes y PDFs:
MEDIA_RESOLUTION_HIGH: 1120 tokensMEDIA_RESOLUTION_MEDIUM: 560 tokensMEDIA_RESOLUTION_LOW: 280 tokensMEDIA_RESOLUTION_UNSPECIFIED(predeterminado): Igual queMEDIA_RESOLUTION_HIGHpara imágenes, yMEDIA_RESOLUTION_MEDIUMpara PDFs.
Para videos, MEDIA_RESOLUTION_LOW y MEDIA_RESOLUTION_MEDIUM corresponden a 70 tokens por fotograma, mientras que MEDIA_RESOLUTION_HIGH enviará 280 tokens por fotograma.
Ten en cuenta que estos son máximos, y el uso real de tokens suele ser ligeramente menor (aproximadamente un 10%).
import os
# Media resolution is only available with `v1alpha`.
client = genai.Client(
api_key=os.environ.get('GEMINI_API_KEY'),
http_options={
'api_version': 'v1alpha',
}
)
# Upload to File API
sample_image = client.files.upload(file="jetpack.png")
media_resolution = 'MEDIA_RESOLUTION_HIGH' # @param ['MEDIA_RESOLUTION_UNSPECIFIED','MEDIA_RESOLUTION_LOW','MEDIA_RESOLUTION_MEDIUM','MEDIA_RESOLUTION_HIGH']
# You can also use types.PartMediaResolutionLevel.MEDIA_RESOLUTION_LOW/MEDIUM/HIGH
count_tokens_response = client.models.count_tokens(
model=GEMINI_3_MODEL_ID,
contents=[
types.Part(
file_data=types.FileData(
file_uri=sample_image.uri,
mime_type=sample_image.mime_type
),
media_resolution=types.PartMediaResolution(
level=media_resolution
),
)
],
)
print(f"The image is worth {count_tokens_response.total_tokens} tokens.")
Firmas de pensamientos
Esta nueva adición no te afectará si estás usando el SDK, ya que es completamente gestionada por los SDKs. Pero si tienes curiosidad, aquí te explicamos lo que sucede detrás de escena.
Si revisas la parte de tu respuesta, notarás una nueva adición: un thought_signature
print(response.parts[1].thought_signature)
Esta firma es utilizada por el modelo cuando quieres tener conversaciones de chat o de múltiples turnos. Ayuda al modelo no solo a recordar lo que se dijo antes, sino también lo que pensó antes o lo que obtuvo de sus herramientas y llamadas a funciones.
Aquí tienes un ejemplo: imagina que le preguntas al modelo la temperatura de hoy. Hará una llamada a una herramienta o usará la búsqueda de Google para obtener el clima y luego te dirá que hará 25 grados. Si luego le preguntas cuál es la humedad, podrá recordar que también obtuvo esa información de la primera llamada y no hará una nueva solicitud.
Más detalles en la documentación.
Migración desde Gemini 2.5
Los modelos Gemini 3 son nuestra familia de modelos más capaz hasta la fecha y ofrecen una mejora progresiva sobre Gemini 2.5 Pro. Al migrar, considera lo siguiente:
- Pensamiento: Si antes usabas ingeniería de prompt compleja (como Chain-of-thought) para forzar a Gemini 2.5 a razonar, prueba Gemini 3 con
thinking_level: "high"y prompts simplificados (más en la guía de pensamiento). - Configuración de temperatura: Si tu código existente establece explícitamente la temperatura (especialmente a valores bajos para salidas deterministas), considera eliminar este parámetro y usar el valor predeterminado de 1.0 de Gemini 3 para evitar posibles problemas de bucle o degradación del rendimiento en tareas complejas.
- Comprensión de PDF y documentos: La resolución OCR predeterminada para PDFs ha cambiado. Si dependías de un comportamiento específico para el análisis de documentos densos, prueba la nueva configuración de
MEDIA_RESOLUTION_HIGHpara asegurar una precisión continua. - Consumo de tokens: La migración a los valores predeterminados de Gemini 3 Pro puede aumentar el uso de tokens para PDFs, pero disminuir el uso de tokens para video. Si las solicitudes ahora exceden la ventana de contexto debido a resoluciones predeterminadas más altas, considera reducir explícitamente la resolución de medios.
- Segmentación de imágenes: Las capacidades de segmentación de imágenes (que devuelven máscaras a nivel de píxel para objetos) no son compatibles con Gemini 3 Pro. Para cargas de trabajo que requieren segmentación de imágenes incorporada, considera seguir utilizando Gemini 3.7 Flash con el pensamiento desactivado (consulta la guía de comprensión espacial
) o Gemini Robotics-ER 1.5
.
Próximos pasos
Referencias útiles de la API:
Consulta el SDK de Google GenAI y su documentación para obtener más detalles sobre el SDK de GenAI.
Ejemplos relacionados
Para ejemplos más detallados usando modelos Gemini, consulta la carpeta Quickstarts del cookbook.
Aprenderás a usar la API en vivo
, a manejar múltiples herramientas
o a usar las habilidades de comprensión espacial
de Gemini.
También deberías revisar todos los modelos gen-media:
- Generación de podcasts y voz usando Gemini TTS
, - Interacción en vivo con Gemini Live
, - Generación de imágenes usando Imagen
, - Generación de video usando Veo
, - Generación de música usando Lyria RealTime
.
Luego, dirígete a la guía de modelos de pensamiento de Gemini
que muestra explícitamente sus resúmenes de pensamientos y puede manejar razonamientos más complejos.
Finalmente, echa un vistazo a la carpeta de ejemplos del cookbook para casos de uso más complejos y demostraciones que mezclan diferentes capacidades.