Lección 33 · 10 min · Gratis

Primeros pasos con agentes gestionados

Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
#     https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

La API de Interacciones proporciona una interfaz unificada para trabajar con modelos y agentes de Gemini. El notebook de Introducción cubre cómo usarla con modelos estándar de Gemini para la generación de texto, conversaciones de múltiples turnos y uso de herramientas.

Este notebook se enfoca en algo diferente: agentes gestionados con el agente antigravity-preview-05-2026.

agent= vs model=

Cuando llamas a la API de Interacciones, eliges entre dos modos:

Parámetro Qué se ejecuta Ideal para
model="gemini-..." Un modelo estándar de Gemini Generación de texto, salida estructurada, llamada a funciones
agent="antigravity-preview-05-2026" Un agente gestionado en un entorno Linux aislado Tareas autónomas: ejecución de código, investigación web, gestión de archivos

Con model=, obtienes una llamada LLM sin estado (consulta el notebook de Introducción). Con agent=, inicias un agente autónomo que puede razonar, planificar, escribir y ejecutar código, navegar por la web y gestionar archivos, todo dentro de un entorno aislado seguro, sin que tú escribas ninguna lógica de orquestación.

Este notebook te guía paso a paso por el modo agente:

  1. Preguntas sencillas: usa el agente como un LLM (funciona, ¡pero es excesivo!)
  2. Conversaciones de múltiples turnos: entorno aislado persistente = memoria integrada
  3. Uso de herramientas: ejecución de código, búsqueda web, operaciones de archivos
  4. Carga de datos en el entorno aislado: inyecta archivos antes de que el agente comience
  5. Creación de agentes personalizados reutilizables: agrupa instrucciones, habilidades y entorno

Configuración

Instalar el SDK

Instala el SDK desde PyPI. Se recomienda usar siempre la última versión.

%pip install -U -q "google-genai>=2.9.0"
     ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 52.7/52.7 kB 1.5 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 818.2/818.2 kB 17.6 MB/s eta 0:00:00
   ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ 246.1/246.1 kB 6.2 MB/s eta 0:00:00
[?25hERROR: pip's dependency resolver does not currently take into account all the packages that are installed. This behaviour is the source of the following dependency conflicts.
google-colab 1.0.0 requires google-auth==2.47.0, but you have google-auth 2.53.0 which is incompatible.
google-cloud-aiplatform 1.148.1 requires google-genai<2.0.0,>=1.66.0; python_version >= "3.10", but you have google-genai 2.4.0 which is incompatible.
google-adk 1.29.0 requires google-genai<2.0.0,>=1.64.0, but you have google-genai 2.4.0 which is incompatible.


Configurar tu clave de API

Para ejecutar la siguiente celda, tu clave de API debe estar almacenada en un Secreto de Colab llamado GEMINI_API_KEY. Si aún no tienes una clave de API o no estás seguro de cómo crear un Secreto de Colab, consulta Autenticación image para ver un ejemplo.

from google.colab import userdata

GEMINI_API_KEY = userdata.get('GEMINI_API_KEY')

Inicializar el cliente del SDK

Con el nuevo SDK, ahora solo necesitas inicializar un cliente con tu clave de API.

import uuid
from google import genai
from google.genai import types
from IPython.display import Markdown

client = genai.Client(api_key=GEMINI_API_KEY)

# The default managed agent.
AGENT = "antigravity-preview-05-2026"

# Generate a unique suffix for this notebook session to prevent agent ID conflicts
UNIQUE_SUFFIX = uuid.uuid4().hex[:8]

print("Client ready!")
Client ready!

1. Preguntas sencillas: el agente como un LLM

La forma más sencilla de usar un agente gestionado es hacerle una pregunta, tal como llamarías a un modelo estándar de Gemini. Pasa agent="antigravity-preview-05-2026" y environment="remote" para crear un nuevo entorno aislado de Linux para el agente.

Esto funciona, pero es un poco como conducir un coche de Fórmula 1 al supermercado: el agente tiene capacidades de ejecución de código, búsqueda web y gestión de archivos que están inactivas para una simple pregunta de hechos.

interaction = client.interactions.create(
    agent=AGENT,
    input="What is the capital of France?",
    environment="remote",
)

Markdown(interaction.output_text)
/tmp/ipykernel_2790/271594579.py:1: UserWarning: Interactions usage is experimental and may change in future versions.
  interaction = client.interactions.create(
<IPython.core.display.Markdown object>
# The response also includes metadata about the agent's sandbox.
print(f"Status:         {interaction.status}")
print(f"Interaction ID: {interaction.id}")
print(f"Environment ID: {interaction.environment_id}")
Status:         completed
Interaction ID: v1_Chc1SlFNYXVmUE51YXNqckVQM09laGlBYxIXNUpRTWF1ZlBOdWFzanJFUDNPZWhpQWM
Environment ID: fc155701-ea68-4846-973f-f2943e35cdde

Observa el environment_id en la respuesta. Ese es el entorno aislado persistente de Linux del agente. Incluso para esta pregunta sencilla, se aprovisionó un contenedor completo. Aprovechemos esa persistencia a continuación.

2. Conversaciones de múltiples turnos

Dado que cada agente se ejecuta en un entorno aislado persistente, puedes continuar donde lo dejaste reutilizando el environment_id y vinculando los turnos con previous_interaction_id.

Esto es fundamentalmente diferente de las llamadas model= sin estado. El agente tiene un verdadero entorno persistente: los archivos que crea permanecen, los paquetes que instala siguen disponibles y el contexto de la conversación se conserva.

# Turn 1: Introduce yourself.
turn1 = client.interactions.create(
    agent=AGENT,
    input="Hi! My name is Alice and I'm a software engineer. Remember that in a knowledge.md doc.",
    environment= "remote",
)

Markdown(f"**Turn 1:** {turn1.output_text}")
<IPython.core.display.Markdown object>
# Turn 2: Ask whether the agent remembers.
# Pass environment_id and previous_interaction_id to continue the conversation.
turn2 = client.interactions.create(
    agent=AGENT,
    input="what's my name and what do I do?",
    environment= turn1.environment_id,
)

Markdown(f"**Turn 2:** {turn2.output_text}")
<IPython.core.display.Markdown object>

El agente recordó entre turnos porque pasaste environment con el ID del entorno anterior, el mismo entorno aislado, lo que significa los mismos archivos.

Podrías haber logrado el mismo resultado usando previous_interaction_id para mantener el historial de la conversación anterior, pero eso no habría mostrado las especificidades del entorno.

Así es como se construyen flujos de trabajo con estado y de múltiples turnos. Consulta el notebook de Introducción para conversaciones de múltiples turnos basadas en model= usando solo previous_interaction_id (sin entornos).

3. Uso de herramientas: donde el agente brilla

Aquí es donde los agentes gestionados van más allá de un modelo de chat estándar. El agente antigravity-preview-05-2026 tiene herramientas integradas que usa de forma autónoma; no las declaras, solo describes tu objetivo y el agente averigua qué usar.

Herramienta Descripción
bash Ejecuta comandos de shell en el entorno aislado
google_search Busca en la web información actual
url_context Obtiene y extrae texto de URLs
write_file Crea o sobrescribe archivos en el entorno aislado
read_file Lee el contenido de los archivos del entorno aislado
list_files Lista el contenido del directorio
delete_file Elimina archivos del entorno aislado

Para las herramientas estándar basadas en model= (fundamentación de Google Search, ejecución de código, llamada a funciones), consulta el notebook de Introducción y los notebooks de herramientas dedicados:

Ejecución de código

Haz una pregunta computacional y el agente escribirá código, lo ejecutará en su entorno aislado y devolverá el resultado verificado.

interaction = client.interactions.create(
    agent=AGENT,
    input=(
        "Write a Python script that computes the first 20 Fibonacci numbers. "
        "Run it and show the output."
    ),
    environment="remote",
)

Markdown(interaction.output_text)
<IPython.core.display.Markdown object>

Inspección de pasos: lo que realmente hizo el agente

El campo steps en la respuesta muestra la cadena de razonamiento del agente: sus pensamientos, llamadas a herramientas, resultados de herramientas y salida final. Esto es útil para depurar y comprender el comportamiento del agente.

# Inspect the steps from the Fibonacci interaction above.
for i, step in enumerate(interaction.steps):
    step_type = step.type
    print(f"--- Step {i} [{step_type}] ---")

    # Tool call steps show which tool was invoked and with what arguments.
    if hasattr(step, "name") and step.name:
        print(f"  Tool: {step.name}")
        if hasattr(step, "arguments"):
            args_str = str(step.arguments)[:300]
            print(f"  Args: {args_str}")

    # Content steps contain the agent's text output.
    if hasattr(step, "content") and step.content:
        for c in step.content:
            if hasattr(c, "text"):
                print(f"  Text: {c.text[:300]}")
    print()
--- Step 0 [thought] ---

--- Step 1 [function_call] ---
  Tool: write_file
  Args: {'path': '/fibonacci.py', 'content': 'def fibonacci(n):\n    if n <= 0:\n        return []\n    elif n == 1:\n        return [0]\n    \n    fib_sequence = [0, 1]\n    while len(fib_sequence) < n:\n        fib_sequence.append(fib_sequence[-1] + fib_sequence[-2])\n    return fib_sequence\n\nif __name_

--- Step 2 [function_result] ---
  Tool: write_file

--- Step 3 [code_execution_call] ---

--- Step 4 [code_execution_result] ---

--- Step 5 [model_output] ---
  Text: I have created and executed a Python script to compute the first 20 Fibonacci numbers.

### Python Script (`fibonacci.py`)

Here is the code written to `fibonacci.py`:

```python
def fibonacci(n):
    if n <= 0:
        return []
    elif n == 1:
        return [0]
    
    fib_sequence = [0, 1]

Búsqueda web

El agente puede buscar en la web de forma autónoma cuando necesita información actualizada.

interaction = client.interactions.create(
    agent=AGENT,
    input="What were the top 3 news stories about Google this week? Summarize them briefly.",
    environment="remote",
)

Markdown(interaction.output_text)
<IPython.core.display.Markdown object>

Operaciones de archivos

El agente puede crear, leer y gestionar archivos en su entorno aislado. Los archivos persisten dentro del entorno entre turnos.

# Ask the agent to create a file, run it, and show results.
interaction = client.interactions.create(
    agent=AGENT,
    input=(
        "Create a Python file called 'analysis.py' that generates 50 random numbers, "
        "computes mean, median, and standard deviation, then prints the results. "
        "Run it and show the output."
    ),
    environment="remote",
)

Markdown(interaction.output_text)
<IPython.core.display.Markdown object>

4. Carga de datos en el entorno aislado del agente

Puedes inyectar archivos en el entorno del agente antes de que comience usando sources. Así es como proporcionas datos, configuración o código para que el agente trabaje.

Tipo de fuente Descripción Ideal para
inline Incrusta contenido directamente (máx. 75 KB) Archivos de configuración, scripts pequeños
gcs Carga desde Google Cloud Storage Grandes conjuntos de datos
repository Carga desde GitHub Repositorios de código
# Inject a CSV file inline and ask the agent to analyze it.
csv_data = """name,age,city,score
              Alice,28,Paris,92
              Bob,35,London,87
              Charlie,42,Berlin,95
              Diana,31,Tokyo,88
              Eve,26,Sydney,91"""

interaction = client.interactions.create(
    agent=AGENT,
    input="Read the file data.csv, analyze it, and tell me who scored the highest.",
    environment={
        "type": "remote",
        "sources": [
            {
                "type": "inline",
                "content": csv_data,
                "target": "/workspace/data.csv",
            }
        ],
    },
)

Markdown(interaction.output_text)
<IPython.core.display.Markdown object>

También puedes cargar desde otras fuentes:

# From Google Cloud Storage
{"type": "gcs", "source": "gs://my-bucket/data/", "target": "/workspace/data/"}

# From a GitHub repository
{"type": "repository", "source": "https://github.com/user/repo", "target": "/workspace/repo/"}

Puedes combinar varias fuentes en una sola solicitud; el agente tendrá acceso a todas ellas al inicio.

Consejo profesional: Puedes usar eso para agregar habilidades a tu agente, como verás a continuación.

5. Creación de agentes personalizados reutilizables

Hasta ahora, cada interacción ha utilizado el agente base antigravity-preview-05-2026 con instrucciones en línea. Una vez que hayas encontrado una configuración que funcione bien, puedes persistirla en un agente personalizado con nombre que agrupe:

  • Instrucciones: prompt del sistema que define el comportamiento del agente
  • Entorno: entorno aislado preconfigurado con archivos y fuentes
  • Habilidades: archivos SKILL.md que enseñan al agente capacidades especializadas

Este es el flujo de trabajo recomendado:

  1. Prototipa con agent="antigravity-preview-05-2026": itera sobre instrucciones, fuentes y prompts
  2. Crea un agente con nombre a través del endpoint /agents
  3. Invoca tu agente por nombre desde cualquier cliente

Creación de un agente personalizado

# Create a custom data analysis agent using the SDK.
my_agent = client.agents.create(
    id=f"my-data-analyst-{UNIQUE_SUFFIX}",
    base_agent=AGENT,
    system_instruction=(
        "You are a data analysis assistant. "
        "Always write Python code using pandas to answer questions. "
        "Show your code and output clearly. "
        "When creating visualizations, save them as PNG files."
    ),
    base_environment={
        "type": "remote",
    },
)

print(f"✓ Agent created: {my_agent.id}")
✓ Agent created: my-data-analyst

Uso de un agente personalizado

Una vez creado, invoca tu agente por su nombre. Seguirá sus instrucciones automáticamente.

# Invoke the custom agent.
interaction = client.interactions.create(
    agent=f"my-data-analyst-{UNIQUE_SUFFIX}",
    input=(
        "Generate a sample dataset of 100 sales records with columns: "
        "product, region, revenue, quantity. "
        "Find the top 5 products by total revenue and show the analysis."
    ),
    environment="remote",
)

Markdown(interaction.output_text)
<IPython.core.display.Markdown object>

Creación de un agente con datos precargados

Puedes definir el entorno del agente con fuentes, de modo que los datos estén listos antes de que el agente comience:

# Create an agent with GCS sources pre-loaded using the SDK.
my_slides_agent = client.agents.create(
    id=f"my-gemini-api-agent-{UNIQUE_SUFFIX}",
    base_agent=AGENT,
    system_instruction=(
        "You are a software engineer speciliazed in the Gemini API. "
        "Use the skills available in /.agents/skills/ to create amazing apps."
    ),
    base_environment={
        "type": "remote",
        "sources": [
            {
                "type": "repository",
                "source": "https://github.com/google-gemini/gemini-skills",
                "target": "/.agents/skills",
            }
        ],
    },
)

print(f"✓ Agent created: {my_slides_agent.id}")
✓ Agent created: my-gemini-api-agent
# Invoke the custom agent.
interaction = client.interactions.create(
    agent=f"my-gemini-api-agent-{UNIQUE_SUFFIX}",
    input="Tell me what you can do with your skills?",
    environment="remote",
)

Markdown(interaction.output_text)
<IPython.core.display.Markdown object>

Bifurcación desde un entorno existente

Si ya has configurado un entorno aislado que te gusta (paquetes instalados, archivos creados, etc.), puedes bifurcarlo en un nuevo agente usando el environment_id de una interacción anterior:

my_forked_agent = client.agents.create(
    id="my-forked-agent",
    base_agent=AGENT,
    system_instruction="Your custom instructions here.",
    base_environment={"env_id": "YOUR_ENVIRONMENT_ID"},
)

Esto captura el estado exacto de ese entorno aislado: todos los paquetes instalados, archivos y configuración.

my_forked_agent = client.agents.create(
    id="my-forked-agent",
    base_agent="my-gemini-api-agent",
    system_instruction="I want all your apps to use the Live API",
    base_environment={"env_id": interaction.environment_id},
)
print(f"✓ Agent forked: {my_forked_agent.id}")
✓ Agent forked: my-forked-agent

Gestión de agentes (CRUD)

El endpoint /agents admite la gestión completa del ciclo de vida:

# List all your agents.
print("Your agents:")
for agent in client.agents.list().agents:
    print(f"- {agent.id}")

# Get a specific agent's details.
agent = client.agents.get(id="my-data-analyst")
print(f"\nAgent details for {agent.id}:")
print(f"Base agent: {agent.base_agent}")
print(f"System instruction: {agent.system_instruction}")
Your agents:
- $AGENT_ID
- code-reviewer
- data-analyst
- my-gemini-api-agent
- my-forked-agent

Agent details for my-data-analyst:
Base agent: antigravity-preview-05-2026
System instruction: You are a data analysis assistant. Always write Python code using pandas to answer questions. Show your code and output clearly. When creating visualizations, save them as PNG files.
# Clean up: delete the agents you created.
for agent_name in ["my-data-analyst", "my-forked-agent", "my-gemini-api-agent"]:
    try:
        client.agents.delete(id=agent_name)
        print(f"✓ Deleted {agent_name}")
    except Exception as e:
        print(f"  Failed to delete {agent_name}: {e}")
✓ Deleted my-data-analyst
✓ Deleted my-forked-agent
✓ Deleted my-gemini-api-agent

Estructura de directorios del agente

Detrás de la API, un agente se define mediante un conjunto simple de archivos. Esto es lo que se implementa cuando creas uno:

my-agent/
├── agent.yaml       # Configuration: base agent, tools, environment
├── AGENTS.md        # System instructions (loaded automatically)
├── skills/          # Custom SKILL.md files that extend capabilities
└── workspace/       # Files seeded into the remote sandbox at startup
  • agent.yaml se asigna directamente al recurso de la API /agents
  • AGENTS.md proporciona instrucciones del sistema, cargadas automáticamente por el arnés
  • skills/ contiene archivos SKILL.md especializados que el agente descubre y usa
  • Los archivos workspace/ se inyectan en el entorno aislado al inicio

Esta estructura basada en archivos hace que los agentes sean fáciles de controlar por versiones, compartir e iterar. Consulta la documentación para obtener más detalles.

6. Streaming

Para tareas más largas, habilita el streaming con stream=True para obtener actualizaciones en tiempo real mientras el agente trabaja. En lugar de esperar la respuesta completa, recibes un flujo de Eventos Enviados por el Servidor (SSE) que te permiten mostrar el progreso al usuario.

Tipos de eventos

El flujo entrega eventos que te informan lo que está haciendo el agente:

Tipo de evento Significado Qué hacer
interaction.created Se creó la interacción Almacena el id para futuras referencias
interaction.status_update El estado cambió (por ejemplo, in_progress) Actualiza el indicador de estado de la interfaz de usuario
step.start Comenzó un nuevo paso (pensamiento, llamada a herramienta, salida) Muestra un indicador de carga
step.delta Contenido incremental: un fragmento de texto, pensamiento o salida de herramienta Añadir a la pantalla: este es el contenido principal
step.stop Un paso completado Oculta el indicador de carga
interaction.completed El agente terminó todo el trabajo Finaliza la interfaz de usuario

Los eventos step.delta son donde reside el contenido. Cada delta tiene un type (por ejemplo, text, thought, function_call, function_result) y contenido que puedes renderizar incrementalmente.

# Stream a response and collect the text as it arrives.
stream = client.interactions.create(
    agent=AGENT,
    input="Write a short poem about the ocean.",
    stream=True,
    environment="remote",
)

collected_text = []

for event in stream:
    # Show the event type so you can see the lifecycle.
    if event.event_type in ("interaction.created", "step.start", "step.stop", "interaction.completed"):
        print(f"[{event.event_type}]")

    # step.delta events carry the actual content.
    elif event.event_type == "step.delta":
        delta = event.delta
        if hasattr(delta, "text") and delta.text:
            print(delta.text, end="", flush=True)
            collected_text.append(delta.text)

print(f"\n\n--- Collected {len(collected_text)} text chunks ---")
[interaction.created]
[step.start]
[step.stop]
[step.start]
Endless cradle of deep, dark blue,
Where whispered secrets of the wind come through.
With restless tides that rise and fall,
You sing an ancient song to all.

A shimmering mirror to the sky,
Where white-winged gulls and currents fly.
In depth and silence, wild and free,
The timeless heart of the mighty sea.[step.stop]
[interaction.completed]


--- Collected 4 text chunks ---

7. Funciones avanzadas

Configuración de red

Por defecto, el entorno aislado del agente tiene acceso de red saliente sin restricciones. Puedes controlar esto con el campo network:

  • Permitir dominios específicos: solo se permiten solicitudes a los dominios listados
  • Inyectar credenciales: agrega automáticamente encabezados (claves de API, tokens) a las solicitudes salientes
  • Deshabilitar red: establece network: "disabled" para bloquear todo el tráfico saliente
# Allow the agent to call only the Gemini API, with an auto-injected API key.
interaction = client.interactions.create(
    agent=AGENT,
    input="Use curl to call the Gemini API and list available models. Show the first 3.",
    environment={
        "type": "remote",
        "network": {
            "allowlist": [
                {
                    "domain": "generativelanguage.googleapis.com",
                    "transform": [{"x-goog-api-key": GEMINI_API_KEY}],
                },
            ]
        },
    },
)

Markdown(interaction.output_text)
<IPython.core.display.Markdown object>

Descargar instantáneas del entorno

Puedes descargar todos los archivos que el agente creó o modificó como un archivo tar. Esto te permite recuperar los productos de trabajo del agente (código, datos, informes) del entorno aislado.

import subprocess
import tarfile
import os

# Create an interaction where the agent produces files.
interaction = client.interactions.create(
    agent=AGENT,
    input=(
        "Create a directory called 'project' with a README.md and a hello.py script. "
        "List the files you created."
    ),
    environment="remote",
)

env_id = interaction.environment_id
print(f"Environment ID: {env_id}")

Markdown(interaction.output_text)
Environment ID: cbcfb191-7b88-46dc-ba0b-008e32ed7cb5
<IPython.core.display.Markdown object>
# Download the environment snapshot.
download_url = (
    f"https://generativelanguage.googleapis.com/v1beta/"
    f"files/environment-{env_id}:download?alt=media"
)

result = subprocess.run(
    ["curl", "-L", "-s", "-o", "snapshot.tar",
     "-H", f"x-goog-api-key: {GEMINI_API_KEY}",
     download_url],
    capture_output=True, text=True,
)

if os.path.exists("snapshot.tar") and os.path.getsize("snapshot.tar") > 0:
    with tarfile.open("snapshot.tar") as tar:
        print("Files in snapshot:")
        for member in tar.getmembers():
            print(f"  {member.name} ({member.size} bytes)")
else:
    print("Snapshot not available (environment may have expired).")
Files in snapshot:
  . (0 bytes)
  ./project (0 bytes)
  ./project/README.md (74 bytes)
  ./project/hello.py (93 bytes)

Próximos pasos

Has recorrido las capacidades principales de los agentes gestionados:

  1. ✅ Preguntas y respuestas sencillas: el agente puede responder preguntas como un LLM
  2. ✅ Múltiples turnos: el entorno aislado persistente permite conversaciones con estado
  3. ✅ Herramientas integradas: ejecución de código, búsqueda web, gestión de archivos
  4. ✅ Carga de datos: inyecta archivos a través de fuentes en línea, GCS o GitHub
  5. ✅ Agentes personalizados: configuraciones reutilizables con instrucciones, habilidades y entorno
  6. ✅ Streaming: actualizaciones en tiempo real mientras el agente trabaja
  7. ✅ Avanzado: control de red e instantáneas del entorno

Más información

Lección del curso «Gemini API Cookbook (quickstarts)» de Google, publicado con licencia Apache 2.0. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Google. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios