Introducción a Deep Research
Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
El agente Deep Research planifica, ejecuta y sintetiza de forma autónoma tareas de investigación de varios pasos en informes detallados y citados. Impulsado por Gemini, navega por complejos paisajes de información (buscando en la web, leyendo páginas, ejecutando código y analizando documentos) para producir informes completos en minutos.
Deep Research es ideal para análisis de mercado, inteligencia competitiva, revisiones de literatura e inmersiones técnicas profundas donde necesitas más que una simple respuesta. Puede consultar más de 100 fuentes en una sola tarea.
Deep Research está disponible exclusivamente a través de la Interactions API y no se puede acceder a él a través de generate_content. Las tareas de investigación se ejecutan de forma asíncrona en segundo plano porque pueden tardar varios minutos en completarse.
La documentación es un buen lugar para empezar a aprender más sobre el agente Deep Research.
Configuración
Configura tu clave de API
Para ejecutar la siguiente celda, tu clave de API debe estar almacenada en un Secreto de Colab llamado GEMINI_API_KEY. Si aún no tienes una clave de API, o no estás seguro de cómo crear un Secreto de Colab, consulta Autenticación
para ver un ejemplo.
from google.colab import userdata
GEMINI_API_KEY=userdata.get('GEMINI_API_KEY')
Instala e inicializa el SDK
%pip install -U -q "google-genai>=2.9.0"
import time
import base64
from google import genai
from IPython.display import display, Image, Markdown
client = genai.Client(api_key=GEMINI_API_KEY)
Selecciona un agente Deep Research
Hay dos versiones del agente Deep Research disponibles:
- Deep Research (
deep-research-preview-04-2026): Optimizado para velocidad y eficiencia con latencia reducida. Ideal para casos de uso interactivos. - Deep Research Max (
deep-research-max-preview-04-2026): Diseñado para la máxima exhaustividad de búsqueda y amplitud del informe. Lo mejor para informes automatizados.
AGENT = "deep-research-preview-04-2026" # @param ["deep-research-preview-04-2026","deep-research-max-preview-04-2026"] {"allow-input":true, isTemplate: true}
A continuación, crea funciones auxiliares para sondear los resultados y mostrar las salidas:
# @title Helper functions (just run this cell)
def wait_for_result(interaction, poll_interval=10):
"""Poll until a background interaction completes or fails."""
print(f"Research started: {interaction.id}")
while True:
result = client.interactions.get(interaction.id)
if result.status == "completed":
return result
elif result.status == "failed":
print(f"Research failed: {result.error}")
return result
print(f" Status: {result.status}", flush=True)
time.sleep(poll_interval)
def display_outputs(result):
"""Display text and image outputs from a completed interaction."""
for step in result.steps:
if step.type == "model_output" and step.content:
for content in step.content:
if content.type == "text":
display(Markdown(content.text))
elif content.type == "image" and content.data:
image_bytes = base64.b64decode(content.data)
display(Image(data=image_bytes))
Ejecuta tu primera tarea de Deep Research
Inicia una tarea de investigación con background=True y sondea el resultado. Deep Research es asíncrono; las tareas pueden tardar varios minutos mientras el agente planifica, busca, lee y sintetiza información.
interaction = client.interactions.create(
input="Research the history of Google TPUs and their impact on AI development.",
agent=AGENT,
background=True,
)
result = wait_for_result(interaction)
display_outputs(result)
Planificación colaborativa
La planificación colaborativa te permite revisar y refinar el plan de investigación antes de que el agente comience su trabajo. Cuando está habilitado, el agente devuelve un plan propuesto en lugar de ejecutarlo inmediatamente. Puedes iterar sobre el plan a través de interacciones de varios turnos.
Paso 1: Solicita un plan
Establece collaborative_planning=True en el agent_config. El agente devuelve un plan de investigación en lugar de un informe completo.
plan_interaction = client.interactions.create(
agent=AGENT,
input="Research Google TPUs vs competitor AI accelerator hardware.",
agent_config={
"type": "deep-research",
"collaborative_planning": True,
},
background=True,
)
result = wait_for_result(plan_interaction)
display_outputs(result)
Paso 2: Refina el plan (opcional)
Usa previous_interaction_id para continuar la conversación e iterar sobre el plan. Mantén collaborative_planning=True para permanecer en modo de planificación. Puedes repetir este paso tantas veces como sea necesario.
refined_plan = client.interactions.create(
agent=AGENT,
input="Add a section comparing power efficiency and total cost of ownership.",
agent_config={
"type": "deep-research",
"collaborative_planning": True,
},
previous_interaction_id=plan_interaction.id,
background=True,
)
result = wait_for_result(refined_plan)
display_outputs(result)
Paso 3: Aprueba y ejecuta
Establece collaborative_planning=False para aprobar el plan e iniciar la investigación completa. El agente ejecutará el plan y devolverá el informe final.
Importante: Debes establecer explícitamente
collaborative_planning=Falseen el turno final. Simplemente enviar "adelante" sin activar la bandera no activará la generación del informe.
final_report = client.interactions.create(
agent=AGENT,
input="Plan looks good!",
agent_config={
"type": "deep-research",
"collaborative_planning": False,
},
previous_interaction_id=refined_plan.id,
background=True,
)
result = wait_for_result(final_report)
display_outputs(result)
Gráficos e infografías integrados
Establece visualization="auto" en el agent_config para habilitar gráficos, diagramas e infografías generados por el agente. Para obtener los mejores resultados, solicita explícitamente elementos visuales en tu consulta; por ejemplo, "Incluye gráficos que muestren tendencias" o "Genera gráficos que comparen la cuota de mercado".
interaction = client.interactions.create(
agent=AGENT,
input="Analyze global semiconductor market trends from 2020 to 2025. Include charts showing market share changes between the top vendors.",
agent_config={
"type": "deep-research",
"visualization": "auto",
},
background=True,
)
result = wait_for_result(interaction)
display_outputs(result)
Consejo: Establecer
visualization="auto"habilita la capacidad, pero el agente genera elementos visuales solo cuando el prompt los solicita. Sé explícito sobre qué gráficos o imágenes quieres.
Streaming en tiempo real
Transmite el progreso de la investigación en tiempo real en lugar de esperar el resultado final. Establece stream=True junto con background=True. Habilita thinking_summaries="auto" para ver los pasos de razonamiento intermedios del agente mientras trabaja.
Tipos de eventos de streaming
| Tipo de evento | Tipo de delta | Descripción |
|---|---|---|
interaction.start |
— | Proporciona el ID de interacción |
content.delta |
thought_summary |
Paso de razonamiento intermedio |
content.delta |
text |
Parte de la salida de texto final |
content.delta |
image |
Una imagen generada (codificada en base64) |
interaction.complete |
— | La investigación ha terminado |
stream = client.interactions.create(
input="Research AI chip market trends. Include charts comparing major vendors.",
agent=AGENT,
background=True,
stream=True,
agent_config={
"type": "deep-research",
"thinking_summaries": "auto",
"visualization": "auto",
},
)
interaction_id = None
last_event_id = None
is_complete = False
def process_stream(stream):
global interaction_id, last_event_id, is_complete
for chunk in stream:
if chunk.event_type == "interaction.start":
interaction_id = chunk.interaction.id
print(f"Interaction started: {interaction_id}")
if chunk.event_id:
last_event_id = chunk.event_id
if chunk.event_type == "content.delta":
if chunk.delta.type == "text":
print(chunk.delta.text, end="", flush=True)
elif chunk.delta.type == "thought_summary":
print(f"\n💭 {chunk.delta.content.text}", flush=True)
elif chunk.delta.type == "image" and chunk.delta.data:
image_bytes = base64.b64decode(chunk.delta.data)
display(Image(data=image_bytes))
elif chunk.event_type in ("interaction.complete", "error"):
is_complete = True
if chunk.event_type == "interaction.complete":
print("\n\n✅ Research Complete")
process_stream(stream)
# Reconnect if the connection drops
while not is_complete and interaction_id:
status = client.interactions.get(interaction_id)
if status.status != "in_progress":
break
print("\n🔄 Reconnecting...")
stream = client.interactions.get(
id=interaction_id, stream=True, last_event_id=last_event_id,
)
process_stream(stream)
Configuración de herramientas
Por defecto, el agente utiliza Google Search, URL Context y Code Execution. Puedes personalizar las herramientas pasando una lista tools. Esto te permite restringir el agente a solo búsqueda web, solo fuentes privadas o una combinación de ambas.
| Herramienta | Valor de tipo | Por defecto | Descripción |
|---|---|---|---|
| Google Search | google_search |
✅ | Busca en la web pública |
| URL Context | url_context |
✅ | Lee y resume páginas web |
| Code Execution | code_execution |
✅ | Ejecuta código para cálculos y análisis de datos |
| MCP Server | mcp_server |
— | Conecta servidores MCP remotos |
| File Search | file_search |
— | Busca en corpus de documentos subidos |
Restringir a solo Google Search
interaction = client.interactions.create(
agent=AGENT,
input="What are the latest developments in quantum computing?",
tools=[{"type": "google_search"}],
background=True,
)
result = wait_for_result(interaction)
display_outputs(result)
Combina varias herramientas
interaction = client.interactions.create(
agent=AGENT,
input="Research the latest breakthroughs in fusion energy. Include data analysis of funding trends.",
tools=[
{"type": "google_search"},
{"type": "url_context"},
{"type": "code_execution"},
],
background=True,
)
result = wait_for_result(interaction)
display_outputs(result)
Servidores MCP remotos
Conecta servidores remotos MCP (Model Context Protocol) para darle al agente acceso a herramientas y servicios externos. Pasa el name, url del servidor y encabezados de autenticación opcionales.
interaction = client.interactions.create(
agent=AGENT,
input="Check the status of my last server deployment.",
tools=[
{
"type": "mcp_server",
"name": "Deployment Tracker",
"url": "https://mcp.example.com/mcp",
"headers": {"Authorization": "Bearer YOUR_TOKEN"},
}
],
background=True,
)
Puedes usar allowed_tools para restringir qué herramientas puede llamar el agente desde un servidor MCP:
interaction = client.interactions.create(
agent=AGENT,
input="Get recent tickets from our issue tracker.",
tools=[
{
"type": "mcp_server",
"name": "Issue Tracker",
"url": "https://mcp.example.com/mcp",
"headers": {"Authorization": "Bearer YOUR_TOKEN"},
"allowed_tools": ["list_tickets", "get_ticket"],
}
],
background=True,
)
Nota: Los servidores MCP admiten no-auth, bearer token y OAuth. Para OAuth, obtén un token con una biblioteca como
google-authy pásalo enheaders.
File Search
Dale al agente acceso a tus propios datos usando File Search. Esto permite que el agente busque en tus corpus de documentos subidos junto con la web.
interaction = client.interactions.create(
input="Compare our 2025 fiscal year report against current public web news.",
agent=AGENT,
background=True,
tools=[
{
"type": "file_search",
"file_search_store_names": ["fileSearchStores/my-store-name"],
}
],
)
result = wait_for_result(interaction)
display_outputs(result)
Entradas multimodales
Deep Research admite entradas multimodales, incluyendo imágenes y documentos (PDF). El agente analiza el contenido proporcionado y realiza una investigación basada en la web contextualizada por las entradas.
Entrada de imagen
Pasa imágenes junto con tu prompt de texto para basar la investigación del agente en contenido visual.
interaction = client.interactions.create(
input=[
{
"type": "text",
"text": "Identify the architectural style shown in this image. Research its historical origins, key characteristics, and notable examples worldwide.",
},
{
"type": "image",
"uri": "https://storage.googleapis.com/generativeai-downloads/images/generated_elephants_giraffes_zebras_sunset.jpg",
},
],
agent=AGENT,
background=True,
)
result = wait_for_result(interaction)
display_outputs(result)
Entrada de documento (PDF)
Pasa documentos directamente como entrada. El agente analiza el documento e investiga su contenido en la web.
interaction = client.interactions.create(
agent=AGENT,
input=[
{"type": "text", "text": "What has been the impact of this research paper? Who are the key authors and what have they worked on since?"},
{
"type": "document",
"uri": "https://arxiv.org/pdf/1706.03762",
"mime_type": "application/pdf",
},
],
background=True,
)
result = wait_for_result(interaction)
display_outputs(result)
Dirigibilidad y formato
Controla la estructura y el tono de la salida proporcionando instrucciones de formato explícitas en tu prompt. Puedes solicitar secciones específicas, tablas de datos o ajustar el tono para diferentes audiencias.
prompt = """
Research the competitive landscape of EV batteries.
Format the output as a technical report with the following structure:
1. Executive Summary (max 200 words)
2. Key Players (Must include a data table comparing capacity, chemistry, and market share)
3. Technology Trends
4. Supply Chain Risks
5. Outlook for 2026-2030
Use a professional, technical tone suitable for an engineering audience.
"""
interaction = client.interactions.create(
input=prompt,
agent=AGENT,
background=True,
)
result = wait_for_result(interaction)
display_outputs(result)
Preguntas de seguimiento
Después de recibir un informe, puedes continuar la conversación usando previous_interaction_id para pedir aclaraciones, resúmenes o elaboraciones sobre secciones específicas sin reiniciar toda la tarea de investigación.
Nota: Las preguntas de seguimiento usan un modelo Gemini estándar (no el agente Deep Research), por lo que se devuelven inmediatamente sin requerir
background=True.
# Use the interaction ID from any completed research task above
follow_up = client.interactions.create(
input="Summarize the key findings in 3 bullet points. What was the most surprising insight?",
model="gemini-3.1-pro-preview",
previous_interaction_id=interaction.id,
)
display(Markdown(follow_up.output_text))
Referencia de configuración del agente
Deep Research usa el parámetro agent_config para controlar el comportamiento. Aquí tienes un resumen de todos los campos configurables:
| Campo | Tipo | Por defecto | Descripción |
|---|---|---|---|
type |
string |
Requerido | Debe ser "deep-research" |
thinking_summaries |
string |
"none" |
"auto" para recibir pasos de razonamiento intermedios durante el streaming |
visualization |
string |
"auto" |
"auto" para habilitar gráficos e imágenes; "off" para deshabilitar |
collaborative_planning |
boolean |
false |
true para habilitar la revisión del plan de varios turnos antes de la investigación |
# Example: all options enabled
interaction = client.interactions.create(
agent=AGENT,
input="Research the competitive landscape of cloud GPUs.",
agent_config={
"type": "deep-research",
"thinking_summaries": "auto",
"visualization": "auto",
"collaborative_planning": False,
},
background=True,
stream=True,
)