Lección 229 · 25 min · Gratis

Guía práctica de modelos OpenAI (parte 1 de 2)

Propósito y audiencia

Este manual sirve como tu guía práctica para seleccionar, usar (prompting) y desplegar el modelo OpenAI adecuado (entre GPT 4.1, o3 y o4-mini) para cargas de trabajo específicas. En lugar de documentación exhaustiva, proporcionamos marcos de decisión accionables y ejemplos del mundo real que ayudan a ingenieros de soluciones, gerentes de cuentas técnicas, arquitectos de socios y profesionales semitécnicos a construir soluciones funcionales rápidamente. El contenido se centra en las capacidades actuales de los modelos, implementaciones específicas por sector y las necesidades actuales de la industria, con rutas claras desde la selección del modelo hasta el despliegue en producción. Cada sección ofrece ejemplos de código concisos y adaptables que puedes aplicar inmediatamente a tus casos de uso, mientras señala recursos existentes para profundizar en temas específicos.

Nota: La siguiente guía prescriptiva y experimentación se ha llevado a cabo con los últimos modelos SOTA (estado del arte) disponibles hoy. Estas métricas están sujetas a cambios en el futuro con diferentes escenarios y plazos en consideración.

Cómo usar este manual

Este manual está organizado en secciones distintas para ayudarte a encontrar rápidamente la información que necesitas. Cada sección cubre un aspecto específico de la selección, implementación y despliegue del modelo.

  1. Propósito y audiencia: Una descripción general de para quién es este manual y qué cubre.
  2. Guía de modelos: Una referencia rápida para ayudarte a seleccionar el modelo adecuado para tus necesidades, incluyendo comparaciones de modelos y diagramas de evolución basados en el mapeo de diferentes escenarios de casos de uso.
  3. Casos de uso:
  4. Del prototipo a la producción: Una lista de verificación para ayudarte a pasar del prototipo a la producción.
  5. Árbol de decisión de adaptación: Un diagrama de flujo para guiar tu selección de modelo basándose en requisitos específicos.
  6. Apéndices: Materiales de referencia que incluyen precios, latencia, patrones de prompt y enlaces a recursos externos.

Para decisiones rápidas, concéntrate en las secciones Guía de modelos y Árbol de decisión de adaptación. Para detalles de implementación, explora los casos de uso específicos relevantes para tus necesidades.

================================================================================

Guía de modelos

2.1 Matriz de introducción a los modelos

Modelo Fortaleza principal Ideal para empezar Advertencias Ruta de escalada / degradación
GPT‑4o Chat de voz / visión en tiempo real Agentes multimodales en vivo Ligeramente por debajo de 4.1 en SOTA (estado del arte) de texto Necesitas razonamiento profundo → o4‑mini
GPT‑4.1 Rey de la precisión de texto de 1 millón de tokens Análisis de documentos largos, revisión de código No puede razonar de forma nativa; costo más alto que los minis Presupuesto ajustado → 4.1‑mini / nano
o3 Agente de uso profundo de herramientas Razonamiento de alto riesgo y múltiples pasos Latencia y precio Costo/latencia → o4‑mini
o4‑mini Razonamiento barato y rápido Lógica de "suficientemente buena" de alto volumen Techo de profundidad vs o3 Precisión crítica → o3

(Tabla completa de precios y utilidad → Sección 6.1)

2.2 Evolución del modelo de un vistazo

La línea de modelos de OpenAI ha evolucionado para abordar necesidades especializadas en diferentes dimensiones. Estos diagramas muestran las familias de modelos actuales y sus relaciones.

Diferencias fundamentales: modelos "o-series" vs "GPT"

OpenAI ofrece dos familias de modelos distintas, cada una con fortalezas únicas:

  • Modelos GPT (4o, 4.1): Optimizados para tareas de propósito general con una excelente capacidad para seguir instrucciones. GPT-4.1 destaca con contextos largos (1 millón de tokens), mientras que GPT-4o tiene variantes para voz en tiempo real, texto a voz y voz a texto. GPT-4.1 también viene en variantes mini y nano, mientras que GPT-4o tiene una variante mini. Estas variantes son más baratas y rápidas que sus contrapartes de tamaño completo.

  • Modelos de la serie o (o3, o4-mini): Especializados en razonamiento profundo y resolución de problemas paso a paso. Estos modelos destacan en tareas complejas de varias etapas que requieren pensamiento lógico y uso de herramientas. Elige estos cuando la precisión y la profundidad del razonamiento sean primordiales. Estos modelos también tienen un parámetro opcional reasoning_effort (que se puede configurar en low, medium o high), que permite a los usuarios controlar la cantidad de tokens utilizados para el razonamiento.

Evolución del modelo OpenAI

OpenAI Model Evolution

Características clave

  • Familia GPT-4.1: Optimizada para el procesamiento de contextos largos con una ventana de contexto de 1 millón de tokens.
  • o3: Especializada en razonamiento profundo de múltiples pasos.
  • o4-mini: Combina capacidades de razonamiento con visión a menor costo.

Cada modelo destaca en diferentes escenarios, con fortalezas complementarias que se pueden combinar para flujos de trabajo complejos.

En este manual solo experimentamos con los modelos de la serie GPT-4.1, o3 y o4-mini. No experimentamos con los modelos de la serie GPT-4o.

================================================================================

3A. Caso de uso: RAG de contexto largo para preguntas y respuestas legales

Long-Context RAG for Legal Q&A

🗂️ Matriz TL;DR

Esta tabla resume las principales opciones tecnológicas y su justificación para esta implementación específica de RAG agéntico de contexto largo.

Capa Elección Utilidad
Fragmentación (Chunking) Divisor consciente de oraciones Divide el documento en 20 fragmentos iguales, respetando los límites de las oraciones.
Enrutamiento gpt-4.1-mini Utiliza la comprensión del lenguaje natural para identificar fragmentos relevantes sin índice de embeddings.
Selección de ruta select(ids=[...]) y scratchpad(text="...") Registra el razonamiento mientras profundiza en la jerarquía del documento.
Citación Nivel de párrafo Equilibra la precisión con el costo; proporciona un contexto significativo para las respuestas.
Síntesis gpt-4.1 (Salida estructurada) Genera respuestas directamente de los párrafos seleccionados con citas.
Verificación o4-mini (LLM como juez) Valida la precisión fáctica y la corrección de las citas.

Nota: Los precios y los identificadores de modelo son precisos a abril de 2025, sujetos a cambios.

Esta sección describe la construcción de un sistema de Generación Aumentada por Recuperación (RAG) diseñado para responder con precisión preguntas sobre textos procesales complejos y extensos, utilizando el Manual de Procedimiento de la Junta de Juicios y Apelaciones de Marcas (TBMP) como caso representativo. El TBMP es un recurso legal esencial que detalla los procedimientos que rigen los litigios de marcas ante la Junta de Juicios y Apelaciones de Marcas de la USPTO, y es consultado frecuentemente por abogados de propiedad intelectual y profesionales legales. Al aprovechar los últimos modelos de OpenAI, el sistema mejora la comprensión y la interpretabilidad del contenido legal denso, permitiendo respuestas precisas y contextualmente conscientes a través de una comprensión avanzada del lenguaje y capacidades de recuperación dinámica.

Estos enfoques también se pueden aplicar a otros casos de uso que requieren una recuperación precisa de información de documentación compleja, como manuales de cumplimiento de atención médica, marcos regulatorios financieros o sistemas de documentación técnica donde la precisión, la citación y la auditabilidad son requisitos de misión crítica.

1. Resumen del escenario

  • Corpus: El documento principal es el Manual de Procedimiento de la Junta de Juicios y Apelaciones de Marcas (TBMP, versión 2024). Este manual contiene reglas y pautas de procedimiento detalladas, con un total de 1194 páginas.
  • Usuarios: Los usuarios objetivo son asociados de litigios de propiedad intelectual (PI) y asistentes legales que necesitan respuestas rápidas y precisas a preguntas de procedimiento basadas únicamente en el TBMP.
  • Preguntas típicas: Los usuarios plantean preguntas que requieren síntesis y citación, como:
    1. "¿Cuáles son los requisitos para presentar una moción para obligar a la divulgación según el TBMP?"
    2. "¿Qué plazos se aplican a las conferencias de descubrimiento según lo especificado en el manual?"
    3. "Explica cómo la Junta maneja las reclamaciones de privilegio abogado-cliente durante las deposiciones según el TBMP."
    4. "Enumera las sanciones de la Regla Federal de Procedimiento Civil 11 que la Junta puede invocar según el TBMP."

Nota: Dependiendo de tu entorno de despliegue específico, es posible que debas adaptar algunos pasos de implementación para que coincidan con los requisitos de tu infraestructura.

Si bien la herramienta de búsqueda de archivos de OpenAI ofrece un buen punto de partida para muchos casos de uso, esta sección presenta un enfoque diferente que aprovecha las ventanas de contexto de un millón de tokens para procesar documentos grandes sin ningún preprocesamiento o base de datos vectorial. El enfoque agéntico descrito aquí permite una ingesta de latencia cero, una granularidad dinámica de recuperación y una trazabilidad de citas de grano fino.

2. Flujo de RAG agéntico

Antes de sumergirnos en la implementación, entendamos el enfoque general:

  1. Carga el documento completo en la ventana de contexto.
  2. Divídelo en 20 fragmentos que respeten los límites de las oraciones.
  3. Pregúntale al modelo qué fragmentos podrían contener información relevante.
  4. Profundiza en los fragmentos seleccionados dividiéndolos aún más.
  5. Repite hasta llegar al contenido a nivel de párrafo.
  6. Genera una respuesta basada en los párrafos seleccionados.
  7. Verifica la respuesta para comprobar su precisión fáctica.

Este enfoque de navegación jerárquica imita cómo un humano podría hojear un documento, centrarse en capítulos relevantes, luego en secciones específicas y finalmente leer solo los párrafos más relevantes.

Hierarchical Router

Sistema RAG agéntico: uso del modelo

Etapa del proceso Modelo utilizado Propósito
Enrutamiento inicial gpt-4.1-mini Identifica qué fragmentos del documento podrían contener información relevante
Navegación jerárquica gpt-4.1-mini Continúa profundizando para encontrar los párrafos más relevantes
Generación de respuestas gpt-4.1 Crea una respuesta estructurada con citas de los párrafos seleccionados
Verificación de respuestas o4-mini Valida la precisión fáctica y el uso adecuado de las citas

Este enfoque de preprocesamiento cero aprovecha las grandes ventanas de contexto para navegar por los documentos sobre la marcha, imitando cómo un humano hojearía un documento para encontrar información relevante.

3. Implementación

Implementemos este enfoque paso a paso.

Comienza instalando los paquetes requeridos.

%pip install tiktoken pypdf nltk openai pydantic --quiet
Note: you may need to restart the kernel to use updated packages.

3.1 Carga de documentos

Primero, carguemos el documento y verifiquemos su tamaño. Para esta guía, nos centraremos en las secciones 100-900, que cubren los aspectos procesales centrales hasta la Revisión de la Decisión de la Junta. Las secciones 1000 y posteriores (Interferencias, Procedimientos de Uso Concurrente, Apelaciones Ex Parte) son procedimientos especializados fuera de nuestro alcance actual.

import requests
from io import BytesIO
from pypdf import PdfReader
import re
import tiktoken
from nltk.tokenize import sent_tokenize
import nltk
from typing import List, Dict, Any

# Download nltk data if not already present
nltk.download('punkt_tab')

def load_document(url: str) -> str:
    """Load a document from a URL and return its text content."""
    print(f"Downloading document from {url}...")
    response = requests.get(url)
    response.raise_for_status()
    pdf_bytes = BytesIO(response.content)
    pdf_reader = PdfReader(pdf_bytes)
    
    full_text = ""
    

    max_page = 920  # Page cutoff before section 1000 (Interferences)
    for i, page in enumerate(pdf_reader.pages):
        if i >= max_page:
            break
        full_text += page.extract_text() + "\n"
    
    # Count words and tokens
    word_count = len(re.findall(r'\b\w+\b', full_text))
    
    tokenizer = tiktoken.get_encoding("o200k_base")
    token_count = len(tokenizer.encode(full_text))
    
    print(f"Document loaded: {len(pdf_reader.pages)} pages, {word_count} words, {token_count} tokens")
    return full_text

# Load the document
tbmp_url = "https://www.uspto.gov/sites/default/files/documents/tbmp-Master-June2024.pdf"
document_text = load_document(tbmp_url)

# Show the first 500 characters
print("\nDocument preview (first 500 chars):")
print("-" * 50)
print(document_text[:500])
print("-" * 50)
[nltk_data] Downloading package punkt_tab to
[nltk_data]     /Users/kmurali/nltk_data...
[nltk_data]   Package punkt_tab is already up-to-date!
Downloading document from https://www.uspto.gov/sites/default/files/documents/tbmp-Master-June2024.pdf...
Document loaded: 1194 pages, 595197 words, 932964 tokens

Document preview (first 500 chars):
--------------------------------------------------
TRADEMARK TRIAL AND
APPEAL BOARD MANUAL
OF PROCEDURE (TBMP)
 June 2024
June   2024
United States Patent and Trademark Office
PREFACE TO THE JUNE 2024 REVISION
The June 2024 revision of the Trademark Trial and Appeal Board Manual of Procedure is an update of the
June 2023 edition. This update is moderate in nature and incorporates relevant case law issued between March
3, 2023 and March 1, 2024.
The title of the manual is abbreviated as “TBMP.” A citation to a section of the manual may be written
--------------------------------------------------

¡Podemos ver que el documento tiene más de 900 mil tokens de largo! Si bien podríamos encajar eso en la longitud de contexto de GPT 4.1, también queremos tener citas verificables, por lo que procederemos con una estrategia de fragmentación recursiva.

3.2 Divisor mejorado de 20 fragmentos con tamaño mínimo de tokens

Ahora, creemos una función mejorada para dividir el documento en 20 fragmentos, asegurando que cada uno tenga un tamaño mínimo de tokens y respetando los límites de las oraciones.

20 es un número elegido empíricamente para este documento/tarea específico y podría necesitar ajustarse para otros documentos según el tamaño y la estructura (cuanto mayor sea el número, más finos serán los fragmentos). Sin embargo, el principio clave aquí es dividir secciones del documento para permitir que el modelo de lenguaje decida los componentes relevantes. Este mismo razonamiento también se aplica al parámetro max_depth que se introducirá más adelante en el manual.

# Global tokenizer name to use consistently throughout the code
TOKENIZER_NAME = "o200k_base"

def split_into_20_chunks(text: str, min_tokens: int = 500) -> List[Dict[str, Any]]:
    """
    Split text into up to 20 chunks, respecting sentence boundaries and ensuring
    each chunk has at least min_tokens (unless it's the last chunk).
    
    Args:
        text: The text to split
        min_tokens: The minimum number of tokens per chunk (default: 500)
    
    Returns:
        A list of dictionaries where each dictionary has:
        - id: The chunk ID (0-19)
        - text: The chunk text content
    """
    # First, split the text into sentences
    sentences = sent_tokenize(text)
    
    # Get tokenizer for counting tokens
    tokenizer = tiktoken.get_encoding(TOKENIZER_NAME)
    
    # Create chunks that respect sentence boundaries and minimum token count
    chunks = []
    current_chunk_sentences = []
    current_chunk_tokens = 0
    
    for sentence in sentences:
        # Count tokens in this sentence
        sentence_tokens = len(tokenizer.encode(sentence))
        
        # If adding this sentence would make the chunk too large AND we already have the minimum tokens,
        # finalize the current chunk and start a new one
        if (current_chunk_tokens + sentence_tokens > min_tokens * 2) and current_chunk_tokens >= min_tokens:
            chunk_text = " ".join(current_chunk_sentences)
            chunks.append({
                "id": len(chunks),  # Integer ID instead of string
                "text": chunk_text
            })
            current_chunk_sentences = [sentence]
            current_chunk_tokens = sentence_tokens
        else:
            # Add this sentence to the current chunk
            current_chunk_sentences.append(sentence)
            current_chunk_tokens += sentence_tokens
    
    # Add the last chunk if there's anything left
    if current_chunk_sentences:
        chunk_text = " ".join(current_chunk_sentences)
        chunks.append({
            "id": len(chunks),  # Integer ID instead of string
            "text": chunk_text
        })
    
    # If we have more than 20 chunks, consolidate them
    if len(chunks) > 20:
        # Recombine all text
        all_text = " ".join(chunk["text"] for chunk in chunks)
        # Re-split into exactly 20 chunks, without minimum token requirement
        sentences = sent_tokenize(all_text)
        sentences_per_chunk = len(sentences) // 20 + (1 if len(sentences) % 20 > 0 else 0)
        
        chunks = []
        for i in range(0, len(sentences), sentences_per_chunk):
            # Get the sentences for this chunk
            chunk_sentences = sentences[i:i+sentences_per_chunk]
            # Join the sentences into a single text
            chunk_text = " ".join(chunk_sentences)
            # Create a chunk object with ID and text
            chunks.append({
                "id": len(chunks),  # Integer ID instead of string
                "text": chunk_text
            })
    
    # Print chunk statistics
    print(f"Split document into {len(chunks)} chunks")
    for i, chunk in enumerate(chunks):
        token_count = len(tokenizer.encode(chunk["text"]))
        print(f"Chunk {i}: {token_count} tokens")
    
    return chunks

# Split the document into 20 chunks with minimum token size
document_chunks = split_into_20_chunks(document_text, min_tokens=500)
Split document into 20 chunks
Chunk 0: 42326 tokens
Chunk 1: 42093 tokens
Chunk 2: 42107 tokens
Chunk 3: 39797 tokens
Chunk 4: 58959 tokens
Chunk 5: 48805 tokens
Chunk 6: 37243 tokens
Chunk 7: 33453 tokens
Chunk 8: 38644 tokens
Chunk 9: 49402 tokens
Chunk 10: 51568 tokens
Chunk 11: 49586 tokens
Chunk 12: 47722 tokens
Chunk 13: 48952 tokens
Chunk 14: 44994 tokens
Chunk 15: 50286 tokens
Chunk 16: 54424 tokens
Chunk 17: 62651 tokens
Chunk 18: 47430 tokens
Chunk 19: 42507 tokens

3.3 Función de enrutador con esquema de herramientas mejorado

Ahora, creemos la función de enrutador que seleccionará los fragmentos relevantes y mantendrá un bloc de notas.

Mantener un bloc de notas permite al modelo rastrear los criterios de decisión y el razonamiento a lo largo del tiempo. Esta implementación utiliza un enfoque de dos pasadas con GPT-4.1-mini: primero requiriendo que el modelo actualice el bloc de notas a través de una llamada a la herramienta (tool_choice="required"), luego solicitando una salida JSON estructurada para la selección de fragmentos. Este enfoque proporciona una mejor visibilidad del proceso de razonamiento del modelo al tiempo que garantiza salidas estructuradas consistentes para el procesamiento posterior.

from openai import OpenAI
import json
from typing import List, Dict, Any

# Initialize OpenAI client
client = OpenAI()

def route_chunks(question: str, chunks: List[Dict[str, Any]], 
                depth: int, scratchpad: str = "") -> Dict[str, Any]:
    """
    Ask the model which chunks contain information relevant to the question.
    Maintains a scratchpad for the model's reasoning.
    Uses structured output for chunk selection and required tool calls for scratchpad.
    
    Args:
        question: The user's question
        chunks: List of chunks to evaluate
        depth: Current depth in the navigation hierarchy
        scratchpad: Current scratchpad content
    
    Returns:
        Dictionary with selected IDs and updated scratchpad
    """
    print(f"\n==== ROUTING AT DEPTH {depth} ====")
    print(f"Evaluating {len(chunks)} chunks for relevance")
    
    # Build system message
    system_message = """You are an expert document navigator. Your task is to:
1. Identify which text chunks might contain information to answer the user's question
2. Record your reasoning in a scratchpad for later reference
3. Choose chunks that are most likely relevant. Be selective, but thorough. Choose as many chunks as you need to answer the question, but avoid selecting too many.

First think carefully about what information would help answer the question, then evaluate each chunk.
"""

    # Build user message with chunks and current scratchpad
    user_message = f"QUESTION: {question}\n\n"
    
    if scratchpad:
        user_message += f"CURRENT SCRATCHPAD:\n{scratchpad}\n\n"
    
    user_message += "TEXT CHUNKS:\n\n"
    
    # Add each chunk to the message
    for chunk in chunks:
        user_message += f"CHUNK {chunk['id']}:\n{chunk['text']}\n\n"
    
    # Define function schema for scratchpad tool calling
    tools = [
        {
            "type": "function",
            "name": "update_scratchpad",
            "description": "Record your reasoning about why certain chunks were selected",
            "strict": True,
            "parameters": {
                "type": "object",
                "properties": {
                    "text": {
                        "type": "string",
                        "description": "Your reasoning about the chunk(s) selection"
                    }
                },
                "required": ["text"],
                "additionalProperties": False
            }
        }
    ]
    
    # Define JSON schema for structured output (selected chunks)
    text_format = {
        "format": {
            "type": "json_schema",
            "name": "selected_chunks",
            "strict": True,
            "schema": {
                "type": "object",
                "properties": {
                    "chunk_ids": {
                        "type": "array",
                        "items": {"type": "integer"},
                        "description": "IDs of the selected chunks that contain information to answer the question"
                    }
                },
                "required": [
                    "chunk_ids"
                ],
                "additionalProperties": False
            }
        }
    }
    
    # First pass: Call the model to update scratchpad (required tool call)
    messages = [
        {"role": "system", "content": system_message},
        {"role": "user", "content": user_message + "\n\nFirst, you must use the update_scratchpad function to record your reasoning."}
    ]
    
    response = client.responses.create(
        model="gpt-4.1-mini",
        input=messages,
        tools=tools,
        tool_choice="required"
    )
    
    # Process the scratchpad tool call
    new_scratchpad = scratchpad
    
    for tool_call in response.output:
        if tool_call.type == "function_call" and tool_call.name == "update_scratchpad":
            args = json.loads(tool_call.arguments)
            scratchpad_entry = f"DEPTH {depth} REASONING:\n{args.get('text', '')}"
            if new_scratchpad:
                new_scratchpad += "\n\n" + scratchpad_entry
            else:
                new_scratchpad = scratchpad_entry
            
            # Add function call and result to messages
            messages.append(tool_call)
            messages.append({
                "type": "function_call_output",
                "call_id": tool_call.call_id,
                "output": "Scratchpad updated successfully."
            })
    
    # Second pass: Get structured output for chunk selection
    messages.append({"role": "user", "content": "Now, select the chunks that could contain information to answer the question. Return a JSON object with the list of chunk IDs."})
    
    response_chunks = client.responses.create(
        model="gpt-4.1-mini",
        input=messages,
        text=text_format
    )
    
    # Extract selected chunk IDs from structured output
    selected_ids = []
    if response_chunks.output_text:
        try:
            # The output_text should already be in JSON format due to the schema
            chunk_data = json.loads(response_chunks.output_text)
            selected_ids = chunk_data.get("chunk_ids", [])
        except json.JSONDecodeError:
            print("Warning: Could not parse structured output as JSON")
    
    # Display results
    print(f"Selected chunks: {', '.join(str(id) for id in selected_ids)}")
    print(f"Updated scratchpad:\n{new_scratchpad}")
    
    return {
        "selected_ids": selected_ids,
        "scratchpad": new_scratchpad
    }

3.4 Función de navegación recursiva

Ahora, creemos la función de navegación recursiva que profundiza en el documento. max_depth es el número máximo de niveles para profundizar (teniendo en cuenta los mínimos de tokens):

def navigate_to_paragraphs(document_text: str, question: str, max_depth: int = 1) -> Dict[str, Any]:
    """
    Navigate through the document hierarchy to find relevant paragraphs.
    
    Args:
        document_text: The full document text
        question: The user's question
        max_depth: Maximum depth to navigate before returning paragraphs (default: 1)
    
    Returns:
        Dictionary with selected paragraphs and final scratchpad
    """
    scratchpad = ""
    
    # Get initial chunks with min 500 tokens
    chunks = split_into_20_chunks(document_text, min_tokens=500)
    
    # Navigator state - track chunk paths to maintain hierarchy
    chunk_paths = {}  # Maps numeric IDs to path strings for display
    for chunk in chunks:
        chunk_paths[chunk["id"]] = str(chunk["id"])
    
    # Navigate through levels until max_depth or until no chunks remain
    for current_depth in range(max_depth + 1):
        # Call router to get relevant chunks
        result = route_chunks(question, chunks, current_depth, scratchpad)
        
        # Update scratchpad
        scratchpad = result["scratchpad"]
        
        # Get selected chunks
        selected_ids = result["selected_ids"]
        selected_chunks = [c for c in chunks if c["id"] in selected_ids]
        
        # If no chunks were selected, return empty result
        if not selected_chunks:
            print("\nNo relevant chunks found.")
            return {"paragraphs": [], "scratchpad": scratchpad}
        
        # If we've reached max_depth, return the selected chunks
        if current_depth == max_depth:
            print(f"\nReturning {len(selected_chunks)} relevant chunks at depth {current_depth}")
            
            # Update display IDs to show hierarchy
            for chunk in selected_chunks:
                chunk["display_id"] = chunk_paths[chunk["id"]]
                
            return {"paragraphs": selected_chunks, "scratchpad": scratchpad}
        
        # Prepare next level by splitting selected chunks further
        next_level_chunks = []
        next_chunk_id = 0  # Counter for new chunks
        
        for chunk in selected_chunks:
            # Split this chunk into smaller pieces
            sub_chunks = split_into_20_chunks(chunk["text"], min_tokens=200)
            
            # Update IDs and maintain path mapping
            for sub_chunk in sub_chunks:
                path = f"{chunk_paths[chunk['id']]}.{sub_chunk['id']}"
                sub_chunk["id"] = next_chunk_id
                chunk_paths[next_chunk_id] = path
                next_level_chunks.append(sub_chunk)
                next_chunk_id += 1
        
        # Update chunks for next iteration
        chunks = next_level_chunks

3.5 Ejecuta la navegación mejorada para una pregunta de ejemplo

Ejecutemos la navegación para una pregunta de ejemplo con nuestro enfoque mejorado:

# Run the navigation for a sample question
question = "What format should a motion to compel discovery be filed in? How should signatures be handled?"
navigation_result = navigate_to_paragraphs(document_text, question, max_depth=2)

# Sample retrieved paragraph
print("\n==== FIRST 3 RETRIEVED PARAGRAPHS ====")
for i, paragraph in enumerate(navigation_result["paragraphs"][:3]):
    display_id = paragraph.get("display_id", str(paragraph["id"]))
    print(f"\nPARAGRAPH {i+1} (ID: {display_id}):")
    print("-" * 40)
    print(paragraph["text"])
    print("-" * 40)
Split document into 20 chunks
Chunk 0: 42326 tokens
Chunk 1: 42093 tokens
Chunk 2: 42107 tokens
Chunk 3: 39797 tokens
Chunk 4: 58959 tokens
Chunk 5: 48805 tokens
Chunk 6: 37243 tokens
Chunk 7: 33453 tokens
Chunk 8: 38644 tokens
Chunk 9: 49402 tokens
Chunk 10: 51568 tokens
Chunk 11: 49586 tokens
Chunk 12: 47722 tokens
Chunk 13: 48952 tokens
Chunk 14: 44994 tokens
Chunk 15: 50286 tokens
Chunk 16: 54424 tokens
Chunk 17: 62651 tokens
Chunk 18: 47430 tokens
Chunk 19: 42507 tokens

==== ROUTING AT DEPTH 0 ====
Evaluating 20 chunks for relevance
Selected chunks: 0, 1, 2, 3, 4, 5, 6, 7, 8
Updated scratchpad:
DEPTH 0 REASONING:
The user wants to know the format requirements for filing a motion to compel discovery and how signatures should be handled for such motions. 

Based on the evaluation of chunks:
- Chunks 0, 1, 2, 3, 4, 5, 6, 7, 8 are highly relevant since they cover general requirements for submissions, motions, signatures, service, and specifically for motions and discovery in TTAB proceedings.
- These chunks contain detailed info about electronic filing (via ESTTA), paper filing exceptions, signature requirements, service requirements, format of submissions (including motions), timing rules, and professionals' responsibilities.
- Additionally, the rules for motions to compel, including required attachments, timing, and certification of good faith efforts to resolve discovery disputes, are specifically outlined.
- Chunks 11-19 mostly cover post-trial and appeal procedures, less directly relevant.

I will select these relevant chunks to provide a thorough answer about how motions to compel discovery should be filed and how signatures on such motions are handled.
Split document into 20 chunks
Chunk 0: 3539 tokens
Chunk 1: 2232 tokens
Chunk 2: 1746 tokens
Chunk 3: 3078 tokens
Chunk 4: 1649 tokens
Chunk 5: 2779 tokens
Chunk 6: 2176 tokens
Chunk 7: 1667 tokens
Chunk 8: 1950 tokens
Chunk 9: 1730 tokens
Chunk 10: 1590 tokens
Chunk 11: 1964 tokens
Chunk 12: 1459 tokens
Chunk 13: 2070 tokens
Chunk 14: 2422 tokens
Chunk 15: 1976 tokens
Chunk 16: 2335 tokens
Chunk 17: 2694 tokens
Chunk 18: 2282 tokens
Chunk 19: 982 tokens
Split document into 20 chunks
Chunk 0: 2880 tokens
Chunk 1: 1323 tokens
Chunk 2: 2088 tokens
Chunk 3: 1493 tokens
Chunk 4: 2466 tokens
Chunk 5: 2563 tokens
Chunk 6: 2981 tokens
Chunk 7: 2723 tokens
Chunk 8: 2264 tokens
Chunk 9: 1900 tokens
Chunk 10: 2134 tokens
Chunk 11: 1778 tokens
Chunk 12: 2484 tokens
Chunk 13: 1922 tokens
Chunk 14: 2237 tokens
Chunk 15: 2044 tokens
Chunk 16: 2097 tokens
Chunk 17: 1326 tokens
Chunk 18: 2427 tokens
Chunk 19: 962 tokens
Split document into 20 chunks
Chunk 0: 2341 tokens
Chunk 1: 1724 tokens
Chunk 2: 2042 tokens
Chunk 3: 3225 tokens
Chunk 4: 1617 tokens
Chunk 5: 2247 tokens
Chunk 6: 1741 tokens
Chunk 7: 1914 tokens
Chunk 8: 2027 tokens
Chunk 9: 2596 tokens
Chunk 10: 2366 tokens
Chunk 11: 2164 tokens
Chunk 12: 2471 tokens
Chunk 13: 1821 tokens
Chunk 14: 1496 tokens
Chunk 15: 1712 tokens
Chunk 16: 1909
… (salida recortada)

¡Los resultados de GPT 4.1-mini muestran la extracción iterativa de componentes relevantes en un documento con el bloc de notas explicando su proceso de pensamiento! En la profundidad 1, el modelo identifica "Reglas detalladas para firmas en presentaciones, incluidas mociones" y "uso de ESTTA, formato de firma requerido, incluidas firmas electrónicas con el método de símbolo '/sig/'" como componentes críticos necesarios para responder la consulta.

En la profundidad 2, el bloc de notas demuestra un juicio sofisticado al aislar con precisión qué fragmentos contienen regulaciones vitales sobre firmas electrónicas (fragmentos 5-12) mientras mantiene la conciencia del contenido ausente, señalando "mociones relacionadas con el descubrimiento... deberían estar en fragmentos desde el 400 en adelante (aunque no son completamente visibles aquí...)".

Este proceso muestra cómo GPT 4.1 imita a un analista legal, profundizando iterativamente en el contenido relevante y explicando su razonamiento a lo largo del camino (lo que facilita la depuración de por qué el modelo seleccionó los fragmentos que hizo).

3.6 Generación de respuestas

Ahora, generemos una respuesta usando GPT-4.1 con los párrafos recuperados.

Aquí hacemos un truco ingenioso en el que construimos dinámicamente una Lista de Literales (lo que obliga a que las respuestas del modelo sean una de las opciones que proporcionamos, en este caso los ID de párrafo). Hay algunas restricciones en el número de opciones que podemos proporcionar, por lo que si tu sistema cita más de 500 documentos, esta solución podría no funcionar. En ese caso, puedes tener un filtro para llegar hasta 500 citas potenciales, o puedes pedirle al modelo que cite el ID exacto en su respuesta, luego posprocesar la respuesta para extraer los ID, y así las citas (por ejemplo, podría decir "... [doc 0.0.12]", y podrías usar alguna expresión regular para extraer la cita).

from typing import List, Dict, Any
from pydantic import BaseModel, field_validator

class LegalAnswer(BaseModel):
    """Structured response format for legal questions"""
    answer: str
    citations: List[str]
    
    @field_validator('citations')
    def validate_citations(cls, citations, info):
        # Access valid_citations from the model_config
        valid_citations = info.data.get('_valid_citations', [])
        if valid_citations:
            for citation in citations:
                if citation not in valid_citations:
                    raise ValueError(f"Invalid citation: {citation}. Must be one of: {valid_citations}")
        return citations

def generate_answer(question: str, paragraphs: List[Dict[str, Any]], 
                   scratchpad: str) -> LegalAnswer:
    """Generate an answer from the retrieved paragraphs."""
    print("\n==== GENERATING ANSWER ====")
    
    # Extract valid citation IDs
    valid_citations = [str(p.get("display_id", str(p["id"]))) for p in paragraphs]
    
    if not paragraphs:
        return LegalAnswer(
            answer="I couldn't find relevant information to answer this question in the document.",
            citations=[],
            _valid_citations=[]
        )
    
    # Prepare context for the model
    context = ""
    for paragraph in paragraphs:
        display_id = paragraph.get("display_id", str(paragraph["id"]))
        context += f"PARAGRAPH {display_id}:\n{paragraph['text']}\n\n"
    
    system_prompt = """You are a legal research assistant answering questions about the 
Trademark Trial and Appeal Board Manual of Procedure (TBMP).

Answer questions based ONLY on the provided paragraphs. Do not rely on any foundation knowledge or external information or extrapolate from the paragraphs.
Cite phrases of the paragraphs that are relevant to the answer. This will help you be more specific and accurate.
Include citations to paragraph IDs for every statement in your answer. Valid citation IDs are: {valid_citations_str}
Keep your answer clear, precise, and professional.
"""
    valid_citations_str = ", ".join(valid_citations)
    
    # Call the model using structured output
    response = client.responses.parse(
        model="gpt-4.1",
        input=[
            {"role": "system", "content": system_prompt.format(valid_citations_str=valid_citations_str)},
            {"role": "user", "content": f"QUESTION: {question}\n\nSCRATCHPAD (Navigation reasoning):\n{scratchpad}\n\nPARAGRAPHS:\n{context}"}
        ],
        text_format=LegalAnswer,
        temperature=0.3
    )
    
    # Add validation information after parsing
    response.output_parsed._valid_citations = valid_citations
    
    print(f"\nAnswer: {response.output_parsed.answer}")
    print(f"Citations: {response.output_parsed.citations}")

    return response.output_parsed

# Generate an answer
answer = generate_answer(question, navigation_result["paragraphs"], 
                       navigation_result["scratchpad"])
==== GENERATING ANSWER ====

Answer: A motion to compel discovery must be filed electronically with the Trademark Trial and Appeal Board (TTAB) through ESTTA, unless ESTTA is unavailable due to technical problems or there are extraordinary circumstances, in which case a paper submission may be permitted with a written explanation ("Documents that relate to proceedings before the Trademark Trial and Appeal Board must be filed electronically with the Board through ESTTA"; "The rules require that all submissions must be made to the Board electronically, currently through ESTTA, subject to certain limited exceptions permitting submissions to be made on paper. Any permitted paper submission must be accompanied by a written explanation showing that ESTTA was unavailable due to technical problems, or that extraordinary circumstances are present, and, where required, a Petition to the Director with the requisite petition fee" 0.0.5.0, 0.0.5.5.7.3).

The motion should include a title describing its nature, such as “Motion to Compel,” and should bear the appropriate proceeding number and caption at the top of the first page ("The document should also include a title describing its nature, e.g., 'Motion to Compel'... should bear at the top of the first page both the application serial number, and the inter partes proceeding number and caption" 0.0.5.4).

Every submission, including a motion to compel discovery, must be signed by the party filing it, or by the party’s attorney or other authorized representative. For electronic filings through ESTTA, a conventional handwritten signature is not required; instead, an electronic signature is used. The signatory must personally enter a combination of letters, numbers, spaces, and/or punctuation marks between two forward slash ('/') symbols (e.g., /John Smith/), and the signatory's name and title or position must appear immediately below or adjacent to the signature ("Documents filed electronically, including through ESTTA, do not require a conventional signature. Electronic signatures pursuant to 37 C.F.R. § 2.193(c) are required for electronic filings. The party or its representative enters a 'symbol' that has been adopted as a signature. The Board will accept any combination of letters, numbers, space and/or punctuation marks as a valid signature if it is placed between two forward slash ('/') symbols"; "The first and last name, and the title or position, of the person who signs a document in connection with a trademark application, registration, or proceeding before the Trademark Trial and Appeal Board must be set forth immediately below or adjacent to the signature" 0.0.5.5.6.2, 0.0.5.5.6.0).

If a document is filed on behalf of a party by the party’s attorney or other authorized representative, it must bear the signature of that attorney or representative, unless the document is one required to be signed personally by the party (0.0.5.5.6.3). If an unsigned or improperly signed document is filed, it will not be refused consideration if a properly signed copy is submitted within the time limit set in the notification of the defect by the Board (0.0.5.5.6.4).

In summary: File the motion to compel discovery electronically via ESTTA, use an electronic signature as described above, and ensure the signatory's name and title are included. If filing on paper is necessary, follow the specific requirements for paper submissions and signatures.
Citations: ['0.0.5.0', '0.0.5.4', '0.0.5.5.6.0', '0.0.5.5.6.2', '0.0.5.5.6.3', '0.0.5.5.6.4', '0.0.5.5.7.3']

GPT 4.1 integra eficazmente las citas a lo largo de su respuesta, manteniendo un flujo claro de información. Cada requisito procesal está vinculado a referencias autorizadas específicas (como "0.0.5.0" y "0.0.5.5.6.2"), creando una respuesta que es tanto informativa como precisamente referenciada.

En lugar de simplemente enumerar las citas al final, las entrelaza directamente en el contenido utilizando notación parentética después de cada requisito clave. Este enfoque transforma una recitación estándar de reglas en un análisis legal bien respaldado donde las declaraciones sobre los procedimientos de presentación de ESTTA, los requisitos de firma electrónica y las excepciones de presentación en papel están respaldadas inmediatamente por sus citas regulatorias correspondientes.

3.7 Verificación de respuestas

Primero, veamos los párrafos citados:

cited_paragraphs = []
for paragraph in navigation_result["paragraphs"]:
    para_id = str(paragraph.get("display_id", str(paragraph["id"])))
    if para_id in answer.citations:
        cited_paragraphs.append(paragraph)
    

# Display the cited paragraphs for the audience
print("\n==== CITED PARAGRAPHS ====")
for i, paragraph in enumerate(cited_paragraphs):
    display_id = paragraph.get("display_id", str(paragraph["id"]))
    print(f"\nPARAGRAPH {i+1} (ID: {display_id}):")
    print("-" * 40)
    print(paragraph["text"])
    print("-" * 40)
==== CITED PARAGRAPHS ====

PARAGRAPH 1 (ID: 0.0.5.0):
----------------------------------------
104  Business to be Conducted in Writing
37 C.F.R. § 2.190(b)  Electronic trademark documents. … Documents that r elate to proceedings before
the Trademark Trial and Appeal Board must be filed electronically with the Board through ESTTA. 37 C.F.R. § 2.191 Action of the Office based on the written record. All business with the Office must be
transacted in writing. The action of the Office will be based exclusively on the written record. No consideration
will be given to any alleged oral promise, stipulation, or understanding when there is disagreement or doubt. With the exceptions of discovery conferences with Board participation, see TBMP § 401.01, and telephone
conferences, see TBMP § 413.01 and TBMP § 502.06, all business with the Board should be transacted in
writing. 37 C.F.R. § 2.191 . The personal attendance of parties or their attorne ys or other authorized
representatives at the offices of the Board is unnecessary , except in the case of a pretrial conference as
provided in 37 C.F.R. § 2.120(j), or upon oral argument at final hearing, if a party so desires, as pro vided
in 37 C.F.R. § 2.129. Decisions of the Board will be based exclusively on the written record before it. [Note
1.] Documents filed in proceedings before the Board must be filed through ESTT A. 37 C.F.R. § 2.190(b). See TBMP § 110.01(a). Board proceedings are conducted in English. If a party intends to rely upon an y submissions that are in a
language other than English, the party should also file a translation of the submissions. If a translation is
not filed, the submissions may not be considered. [Note 2.] NOTES:
1. Cf.
----------------------------------------

PARAGRAPH 2 (ID: 0.0.5.4):
----------------------------------------
The document should
also include a title describing its nature, e.g., “Notice of Opposition,” “Answer,” “Motion to Compel,” “Brief
in Opposition to Respondent’s Motion for Summary Judgment,” or “Notice of Reliance.”
Documents filed in an application which is the subject of an inter partes proceeding before the Board should
be filed with the Board, not the Trademark Operation, and should bear at the top of the first page both the
application serial number, and the inter partes proceeding number and caption. Similarly , requests under
Trademark Act § 7, 15 U.S.C. § 1057, to amend, correct, or surrender a registration which is the subject of
a Board inter partes proceeding, and any new power of attorney, designation of domestic representative, or
change of address submitted in connection with such a registration, should be filed with the Board, not with
the Trademark Operation, and should bear at the top of its first page the re gistration number, and the inter
partes proceeding number and the proceeding caption. [Note 2.] 100-14June   2024
TRADEMARK TRIAL AND APPEAL BOARD MANUAL OF PROCEDURE§ 105
NOTES:
1. 37 C.F.R. § 2.194. 2. 37 C.F.R. § 2.194. 106.02  Signature of Submissions
37 C.F.R. § 2.119(e) Every submission filed in an inter partes proceeding, and every request for an extension
of time to file an opposition, must be signed by the party filing it, or by the party’s attorney or other authorized
representative, but an unsigned submission will not be r efused consideration if a signed copy is submitted
to the Office within the time limit set in the notification of this defect by the Office. 37 C.F.R. § 11.14(e) Appearance.
----------------------------------------

PARAGRAPH 3 (ID: 0.0.5.5.6.0):
----------------------------------------
The Office will accept an electronic signature that meets the
requirements of paragraph (c) of this section on correspondence filed on paper or through TEAS or ESTTA. (b)   Copy of original signature. If a copy of an original signature is filed, the filer should retain the
original as evidence of authenticity. If a question of authenticity arises, the Office may require submission
of the original. (c)   Requirements for electronic signature. A person signing a document electronically must:
(1)   Personally enter any combination of letters, numbers, spaces and/or punctuation marks that the
signer has adopted as a signature, placed between two forward slash (“/”) symbols in the signature block
on the electronic submission; or
(2)   Sign the verified statement using some other form of electronic signature specified by the Director. (d)   Signatory must be identified. The first and last name, and the title or position, of the person who
signs a document in connection with a trademark application, registration, or proceeding before the
Trademark Trial and Appeal Board must be set forth immediately below or adjacent to the signature. (e)   Proper person to sign. Documents filed in connection with a trademark application or registration
must be signed as specified in paragraphs (e)(1) through (9) of this section. (2)   Responses, amendments to applications, requests for express abandonment, requests for
reconsideration of final actions, and requests to divide. Responses to Office actions, amendments to
applications, requests for express abandonment, requests for reconsideration of final actions, and requests
to divide must be signed by the owner of the application or registration, someone with legal authority to
bind the owner (e.g.
----------------------------------------

PARAGRAPH 4 (ID: 0.0.5.5.6.2):
----------------------------------------
* * * *
(i)   Certified documents required by statute. When a statute requires that a document be certified, a
copy or facsimile transmission of the certification is not acceptable. Every document filed in an inter partes or e x parte proceeding before the Board, and e very request for an
extension of time to file an opposition, must be signed by the party filing it, or by the party’ s attorney or
other authorized representative, as appropriate, and the signatory must be identified. [Note 1.] Documents filed electronically, including through ESTTA, do not require a conventional signature. Electronic
signatures pursuant to 37 C.F.R. § 2.193(c) are required for electronic filings. The party or its representative
enters a “symbol” that has been adopted as a signature. The Board will accept any combination of letters,
numbers, space and/or punctuation marks as a valid signature if it is placed between two forward slash (“/”)
symbols. [Note 2.] The electronic signature entered on the ESTTA form is sufficient as the required signature
for the entire submission, including in the absence of a signature on any attachment to the filing form. [Note
3.] The electronic filing cover sheet in ESTTA must be signed by the party filing it, the party’s attorney or
other authorized representative, as appropriate. For further information regarding the filing of submissions
using ESTTA, see TBMP § 110. A party may act in its own behalf in a proceeding before the Board, if the party is domiciled in the United
States, or an attorney may represent the party. [Note 4.] See TBMP § 114 (Representation of a Party). When an individual who is a party to a Board proceeding elects to act in the indi vidual's own behalf, the
individual must sign any documents that are filed with the Board.
----------------------------------------

PARAGRAPH 5 (ID: 0.0.5.5.6.3):
----------------------------------------
If a party which is a partnership elects to
act in its own behalf, a partner should sign documents filed by the partnership. If a party which is a corporation
or association elects to act in its own behalf, an officer thereof who is authorized to sign for the corporation
or association should sign for that corporation or association. If joint applicants elect to act on their o wn
behalf, all joint applicants must sign any documents filed with the Board. [Note 5.] If a document is filed on behalf of a party by the party’s attorney or other authorized representative, it must
bear the signature of, and be personally signed or inserted by , that attorney or other representative, unless
June   2024100-17
§ 106.02GENERAL INFORMATION
it is a document required to be signed personally by the party. An attorney or other authorized representative
who signs a document, and then files it with the Board on behalf of a party , should remember that the
signature to the document constitutes a certification of the elements specified in 37 C.F.R. § 11.18(b), and
that a violation of the pro visions of that rule by may result in sanctions or disciplinary action. [Note 6.] SeeTBMP § 114.04 (regarding meaning of the designation “other authorized representati ve”) and TBMP
§ 527.02 (regarding motions for Fed. R. Civ. P. 11 sanctions). A person transmitting paper documents, when
permitted, for filing with the Board may sign a co ver letter or transmittal letter , and the Office does not
require the party, attorney, or authorized representative to sign a cover or transmittal letter. It is not appropriate for one person to sign a document for another person, as, for example, “John Smith, for
John Doe” or “John Doe, by John Smith.” [Note 7.]
----------------------------------------

PARAGRAPH 6 (ID: 0.0.5.5.6.4):
----------------------------------------
A document filed in a proceeding before the Board should include the first and last name, in typed or printed
form, of the person who signed [Note 8]; a description of the capacity in which the person signed (e.g., as
the individual who is a party, if the filing party is an individual; as a corporate officer, if the filing party is
a corporation; or as the filing party’s attorney); and the business address and telephone number of the person. The inclusion of the signing person’s address and phone number on the submission itself is vital in the rare
case any paper or physical submissions permitted under the rules because mail physically sent to the Office
is opened in the Mail Room, and ordinarily the en velopes are discarded there before the mail is sent on to
its ultimate destination within the Office. Thus, the Board rarely sees the return addresses on the mailing
envelopes of papers filed in Board proceedings. In accordance with 37 C.F.R. § 2.193(b), a legible copy of the signed document is to be filed with the Board
because filings are required to be submitted using ESTT A. The original should be retained as e vidence of
authenticity. If a question as to the authenticity of a filed copy arises, the Office may require submission of
the original. [Note 9.] Notwithstanding the requirement that a document filed before the Board be signed, an unsigned document
filed in paper form, when permitted, will not be refused consideration if a signed cop y is submitted to the
Board within the time limit set in the notification of this defect by the Board. [Note 10.] Similarly , an
improperly signed document, whether filed in ESTT A or on paper , when permitted, will not be refused
consideration if a properly signed cop y is submitted to the Board within the time set in the notification of
this defect by the Board.
----------------------------------------

PARAGRAPH 7 (ID: 0.0.5.5.7.3):
----------------------------------------
long, and contain no tabs or other such devices extending beyond the edges of the paper;
(3)   If a paper submission contains dividers, the dividers must not have any extruding tabs or other
devices, and must be on the same size and weight paper as the submission;
(4)   A paper submission must not be stapled or bound;
(5)   All pages of a paper submission must be numbered and exhibits shall be identified in the manner
prescribed in § 2.123(g)(2);
June   2024100-19
§ 106.03GENERAL INFORMATION
(6)   Exhibits pertaining to a paper submission must be filed on paper and comply with the requirements
for a paper submission. (c)   To be handled as confidential, submissions to the Trademark Trial and Appeal Board that are
confidential in whole or part pursuant to § 2.125(f) must be submitted using the “Confidential” selection
available in ESTTA or, where appropriate, under a separate paper cover. Both the submission and its cover
must be marked confidential and must identify the case number and the parties. A copy of the submission
for public viewing with the confidential portions redacted must be submitted concurrently. The rules require that all submissions must be made to the Board electronically, currently through ESTTA,
subject to certain limited e xceptions permitting submissions to be made on paper . Any permitted paper
submission must be accompanied by a written e xplanation showing that ESTTA was unavailable due to
technical problems, or that extraordinary circumstances are present, and, where required, a Petition to the
Director with the requisite petition fee. [Note 1.]
----------------------------------------

El truco de la "Lista de Literales" obliga al modelo a citar solo ID de párrafo específicos (como "0.0.5.4") en lugar de inventar sus propias referencias o resaltar texto aleatorio; imagínalo como crear una "tabla de contenido" digital de la que GPT-4.1 solo puede seleccionar. Esta solución garantiza que obtengas rastros de citas verificables hasta el material fuente exacto, resolviendo un problema importante en RAG de contexto largo.

Finalmente, verifiquemos la respuesta con un enfoque de LLM como juez.

from typing import List, Dict, Any, Literal
from pydantic import BaseModel

class VerificationResult(BaseModel):
    """Verification result format"""
    is_accurate: bool
    explanation: str
    confidence: Literal["high", "medium", "low"]

def verify_answer(question: str, answer: LegalAnswer, 
                 cited_paragraphs: List[Dict[str, Any]]) -> VerificationResult:
    """
    Verify if the answer is grounded in the cited paragraphs.
    
    Args:
        question: The user's question
        answer: The generated answer
        cited_paragraphs: Paragraphs cited in the answer
        
    Returns:
        Verification result with accuracy assessment, explanation, and confidence level
    """
    print("\n==== VERIFYING ANSWER ====")
    
    # Prepare context with the cited paragraphs
    context = ""
    for paragraph in cited_paragraphs:
        display_id = paragraph.get("display_id", str(paragraph["id"]))
        context += f"PARAGRAPH {display_id}:\n{paragraph['text']}\n\n"
    
    # Prepare system prompt
    system_prompt = """You are a fact-checker for legal information.
Your job is to verify if the provided answer:
1. Is factually accurate according to the source paragraphs
2. Uses citations correctly

Be critical and look for any factual errors or unsupported claims.
Assign a confidence level based on how directly the paragraphs answer the question:
- high: The answer is comprehensive, accurate, and directly supported by the paragraphs
- medium: The answer is mostly accurate but may be incomplete or have minor issues
- low: The answer has significant gaps, inaccuracies, or is poorly supported by the paragraphs
"""
    
    response = client.responses.parse(
        model="o4-mini",
        input=[
            {"role": "system", "content": system_prompt},
            {"role": "user", "content": f"""
QUESTION: {question}

ANSWER TO VERIFY:
{answer.answer}

CITATIONS USED: {', '.join(answer.citations)}

SOURCE PARAGRAPHS:
{context}

Is this answer accurate and properly supported by the source paragraphs?
Assign a confidence level (high, medium, or low) based on completeness and accuracy.
            """}
        ],
        text_format=VerificationResult
    )
    
    # Log and return the verification result
    print(f"\nAccuracy verification: {'PASSED' if response.output_parsed.is_accurate else 'FAILED'}")
    print(f"Confidence: {response.output_parsed.confidence}")
    print(f"Explanation: {response.output_parsed.explanation}")
    
    return response.output_parsed

# Verify the answer using only the cited paragraphs
verification = verify_answer(question, answer, cited_paragraphs)

# Display final result with verification
print("\n==== FINAL VERIFIED ANSWER ====")
print(f"Verification: {'PASSED' if verification.is_accurate else 'FAILED'} | Confidence: {verification.confidence}")
print("\nAnswer:")
print(answer.answer)
print("\nCitations:")
for citation in answer.citations:
    print(f"- {citation}")
==== VERIFYING ANSWER ====

Accuracy verification: PASSED
Confidence: high
Explanation: The answer correctly states that motions to compel discovery must be filed electronically through ESTTA, with paper submissions permitted only under the limited exceptions of technical failure or extraordinary circumstances (37 C.F.R. § 2.190(b) and 2.193(b)). It accurately describes the required title and caption placement (TBMP § 105), and it appropriately summarizes the signature requirements for electronic filings (37 C.F.R. § 2.193(c) and TBMP §§ 106.02, 106.02(b)–(e)), including the use of slash‐enclosed electronic signatures and identification of the signatory’s name and title. It also correctly notes the rule regarding defective signatures (37 C.F.R. § 2.119(e) and TBMP § 106.02). The citations align with the source paragraphs. 

==== FINAL VERIFIED ANSWER ====
Verification: PASSED | Confidence: high

Answer:
A motion to compel discovery must be filed electronically with the Trademark Trial and Appeal Board (TTAB) through ESTTA, unless ESTTA is unavailable due to technical problems or there are extraordinary circumstances, in which case a paper submission may be permitted with a written explanation ("Documents that relate to proceedings before the Trademark Trial and Appeal Board must be filed electronically with the Board through ESTTA"; "The rules require that all submissions must be made to the Board electronically, currently through ESTTA, subject to certain limited exceptions permitting submissions to be made on paper. Any permitted paper submission must be accompanied by a written explanation showing that ESTTA was unavailable due to technical problems, or that extraordinary circumstances are present, and, where required, a Petition to the Director with the requisite petition fee" 0.0.5.0, 0.0.5.5.7.3).

The motion should include a title describing its nature, such as “Motion to Compel,” and should bear the appropriate proceeding number and caption at the top of the first page ("The document should also include a title describing its nature, e.g., 'Motion to Compel'... should bear at the top of the first page both the application serial number, and the inter partes proceeding number and caption" 0.0.5.4).

Every submission, including a motion to compel discovery, must be signed by the party filing it, or by the party’s attorney or other authorized representative. For electronic filings through ESTTA, a conventional handwritten signature is not required; instead, an electronic signature is used. The signatory must personally enter a combination of letters, numbers, spaces, and/or punctuation marks between two forward slash ('/') symbols (e.g., /John Smith/), and the signatory's name and title or position must appear immediately below or adjacent to the signature ("Documents filed electronically, including through ESTTA, do not require a conventional signature. Electronic signatures pursuant to 37 C.F.R. § 2.193(c) are required for electronic filings. The party or its representative enters a 'symbol' that has been adopted as a signature. The Board will accept any combination of letters, numbers, space and/or punctuation marks as a valid signature if it is placed between two forward slash ('/') symbols"; "The first and last name, and the title or position, of the person who signs a document in connection with a trademark application, registration, or proceeding before the Trademark Trial and Appeal Board must be set forth immediately below or adjacent to the signature" 0.0.5.5.6.2, 0.0.5.5.6.0).

If a document is filed on behalf of a party by the party’s attorney or other authorized representative, it must bear the signature of that attorney or representative, unless the document is one required to be signed personally by the party (0.0.5.5.6.3). If an unsigned or improperly signed document is filed, it will not be refused consideration if a properly signed copy is submitted within the time limit set in the notification of the defect by the Board (0.0.5.5.6.4).

In summary: File the motion to compel discovery electronically via ESTTA, use an electronic signature as described above, and ensure the signatory's name and title are included. If filing on paper is necessary, follow the specific requirements for paper submissions and signatures.

Citations:
- 0.0.5.0
- 0.0.5.4
- 0.0.5.5.6.0
- 0.0.5.5.6.2
- 0.0.5.5.6.3
- 0.0.5.5.6.4
- 0.0.5.5.7.3

El paso de verificación produce una evaluación limpia y estructurada que hace referencia a regulaciones específicas y verifica metódicamente tanto la precisión de la respuesta como el uso adecuado de las citas. En lugar de simplemente decir "correcto", ofrece un contexto útil al explicar exactamente por qué la respuesta fue correcta, dándote la confianza para luego presentar la respuesta al usuario con citas específicas.

4. Costos de infraestructura

Desglosemos la estructura de costos para este enfoque RAG agéntico:

Costos fijos vs. variables estimados

  • Costos fijos estimados (únicos):

    • RAG tradicional: ~$0.43 (generación de embeddings + metadatos)
    • RAG agéntico: $0.00 (no se requiere preprocesamiento)
  • Costos variables estimados (por consulta):

    • Modelo de enrutador (gpt-4.1-mini):
      • Enrutamiento inicial (20 fragmentos): ~$0.10
      • Dos niveles recursivos: ~$0.20
    • Síntesis (gpt-4.1): ~$0.05
    • Verificación (o4-mini): ~$0.01
    • Total por consulta: ~$0.36

Si bien el costo por consulta es más alto que el RAG tradicional, este enfoque ofrece:

  • Resultados inmediatos en documentos nuevos
  • Citas más precisas
  • Mejor manejo de paráfrasis y preguntas conceptuales
  • Sin sobrecarga de mantenimiento de infraestructura

El costo se puede optimizar mediante:

  • Almacenamiento en caché de resultados para consultas comunes
  • Limitación de tokens máximos en las llamadas al modelo
  • Uso de un enfoque híbrido que prefiltre el documento primero

5. Beneficios y desventajas frente al RAG tradicional

Beneficios

  • Latencia de ingesta cero: Responde preguntas de documentos nuevos inmediatamente, sin preprocesamiento.
  • Navegación dinámica: Imita los patrones de lectura humana al centrarse en secciones prometedoras.
  • Razonamiento transversal: El modelo puede encontrar conexiones entre secciones del documento que podrían pasarse por alto con la recuperación de fragmentos independientes, lo que potencialmente aumenta la precisión de las respuestas generadas y ahorra tiempo en la optimización de las tuberías de recuperación.

Desventajas

  • Mayor costo por consulta: Requiere más computación para cada pregunta en comparación con la recuperación basada en embeddings.
  • Mayor latencia: La navegación jerárquica tarda más en procesarse que las simples búsquedas vectoriales.
  • Escalabilidad limitada: Puede tener dificultades con colecciones de documentos extremadamente grandes donde el preprocesamiento se vuelve más eficiente.

6. Próximos pasos

Hay algunas modificaciones que podemos hacer al enfoque adoptado:

  • Generación de un grafo de conocimiento: Podemos usar la gran ventana de contexto de GPT 4.1-mini para generar iterativamente un grafo de conocimiento detallado, y luego GPT 4.1 puede recorrer este grafo para responder preguntas. De esta manera, solo necesitamos "ingerir" el documento una vez, independientemente de la pregunta.
  • Herramienta de bloc de notas mejorada: La herramienta de bloc de notas podría tener más opciones, como editar o eliminar la memoria pasada. Esto permitiría al modelo elegir lo que sea más relevante para la pregunta en cuestión.
  • Ajustar la profundidad: Podemos ajustar la profundidad de la navegación jerárquica para encontrar el equilibrio adecuado entre costo y rendimiento. Ciertos casos de uso requerirán citas a nivel de oración (como documentos legales), mientras que otros solo requerirán citas a nivel de párrafo (como artículos de noticias).

7. Conclusiones

  1. La ventana de contexto es un superpoder: Las ventanas de contexto de un millón de tokens hacen posible navegar por documentos sobre la marcha.
  2. El enfoque jerárquico imita la lectura humana: El enrutamiento agéntico funciona como un humano hojeando un documento en busca de secciones relevantes.
  3. El bloc de notas permite el razonamiento en varios pasos: Mantener un registro de razonamiento mejora la calidad de la navegación.
  4. Implementación rápida, sin base de datos: Todo el sistema se puede construir solo con llamadas a la API, sin necesidad de infraestructura.
  5. La verificación mejora la fiabilidad: El patrón de LLM como juez detecta errores antes de que lleguen a los usuarios.

================================================================================

3B. Caso de uso: Co-científico de IA para I+D farmacéutica

AI Co-Scientist for Pharma R&D

Esta sección detalla cómo construir un sistema de IA que funcione como un "co-científico" para acelerar el diseño experimental en I+D farmacéutica, centrándose en optimizar un proceso de síntesis de fármacos bajo restricciones específicas.

🗂️ Matriz TL;DR

Esta tabla resume las elecciones tecnológicas centrales y su justificación para esta implementación específica del Co-Científico de IA.

Capa Elección Utilidad
Ideación o4-mini (Agentes de juego de roles paralelos) Genera hipótesis y protocolos diversos de forma rápida y rentable; el juego de roles mejora la creatividad.
Fundamentación Llamadas a herramientas externas (chem_lookup, cost_estimator, outcome_db, etc.) Asegura que los planes se basen en datos del mundo real (propiedades químicas, costos, resultados anteriores).
Clasificación o4-mini (Comparación de torneos por pares) Evaluación matizada más allá de la puntuación simple; selecciona candidatos prometedores de manera eficiente.
Crítica/Síntesis o3 (Revisión y síntesis profunda) Proporciona un análisis riguroso de nivel sénior, identifica riesgos y asegura la validez científica.
Seguridad (Opc.) gpt-4.1-mini (Verificación dirigida) Añade una capa extra de revisión de seguridad especializada antes de la entrega a un humano.
Aprendizaje o3 + Intérprete de código (Análisis de resultados → DB) Captura los resultados experimentales de forma sistemática, permitiendo una mejora continua con el tiempo.
Técnica central Colaboración y escalada multiagente Aprovecha las fortalezas de diferentes modelos (velocidad vs. profundidad) para una tarea de razonamiento compleja y de múltiples pasos.

Nota: Los identificadores de modelo son precisos a abril de 2025, sujetos a cambios.

1. Resumen del escenario

  • Espacio del problema: Optimizar procedimientos experimentales complejos en I+D farmacéutica, como mejorar el rendimiento de síntesis de un nuevo compuesto farmacológico ("XYZ-13") mientras se adhieren a estrictas restricciones.
  • Usuarios: Científicos de investigación y técnicos de laboratorio involucrados en el descubrimiento y desarrollo de fármacos.
  • Solicitudes típicas:
    1. Sugiere 3 protocolos distintos para aumentar el rendimiento de XYZ-13 en ≥15% probando diferentes catalizadores, manteniéndose por debajo de $15k usando reactivos aprobados.
    2. Propón protocolos para optimizar el rendimiento de XYZ-13 por debajo de 60°C (debido a problemas de calor anteriores), explorando diferentes solventes aprobados dentro del presupuesto.
    3. Diseña dos estrategias de rendimiento de XYZ-13 (apuntando a ≥15%): a. una que maximice el rendimiento potencial dentro del presupuesto de $15k, b. una que priorice el costo por debajo de $10k.
  • Restricciones:
    • Presupuestarias: Operar dentro de límites financieros definidos (por ejemplo, $15,000 por serie de experimentos).
    • Regulatorias/Seguridad: Usar solo productos químicos/reactivos preaprobados y adherirse rigurosamente a los protocolos de seguridad.
    • Supervisión humana: Los planes experimentales finales deben ser revisados y validados por un experto humano antes de su ejecución.

Tradicionalmente, optimizar tales experimentos implica semanas de planificación manual, revisión de literatura, trabajo de laboratorio iterativo y análisis. Este enfoque de Co-Científico de IA tiene como objetivo reducir drásticamente el tiempo del ciclo al automatizar la generación de hipótesis, el diseño de protocolos y la evaluación preliminar, lo que permite a los científicos centrarse en la estrategia de nivel superior y la validación final. Cambia el rol del científico de la ejecución manual de los pasos de planificación a la supervisión experta y la colaboración con la IA.

2. Arquitectura (Razonamiento multiagente)

El sistema emplea una arquitectura multiagente que emula un equipo científico de alto rendimiento. Diferentes componentes de IA, actuando en roles especializados (como ideación, crítica y aprendizaje de resultados), colaboran utilizando varios modelos y herramientas para ejecutar el flujo de trabajo.

AI Co-Scientist Architecture

2.1. Entrada y restricciones del científico:

El proceso comienza con el científico definiendo el objetivo, el compuesto objetivo y las restricciones.

from openai import OpenAI
from agent_utils import Context, call_openai, log_json

# Example Initial Input
user_input = {
    "compound": "XYZ-13",
    "goal": "Improve synthesis yield by 15%",
    "budget": 15000,
    "time_h": 48,
    "previous": "Prior attempts failed at high temp; explore potential catalyst effects."
}
ctx = Context(client=OpenAI(), **user_input)

2.2. Ideación (o4-mini + Herramientas):

Múltiples instancias de o4-mini, con prompts de diferentes roles (por ejemplo, Hypothesis Agent, Protocol Agent, Resource Agent), generan planes experimentales en paralelo. Asignar personas distintas fomenta diversas perspectivas y cubre diferentes aspectos del problema simultáneamente durante la fase de ideación.

ROLE_FOCUS = {
    # Hypothesis Agent Prompt
    "hypothesis_agent": """You are a pharmaceutical hypothesis specialist. 
        Focus exclusively on analyzing the compound structure and research goals to generate testable hypotheses. 
        Consider mechanism of action, binding affinity predictions, and potential off-target effects.""",

    # Protocol Agent Prompt
    "protocol_agent"  : """You are a laboratory protocol specialist. 
        Design experimental procedures that will effectively test the provided hypothesis. 
        Focus on experimental conditions, controls, and measurement techniques.""",

    # Resource Agent Prompt
    "resource_agent"  : """You are a laboratory resource optimization specialist. 
        Review the proposed protocol and optimize for efficiency. 
        Identify opportunities to reduce reagent use, equipment time, and overall costs while maintaining scientific validity.""",
}

# Create a structured prompt template for ideation
IDEATION_PROMPT = """You are a pharmaceutical {role} specialist. Your goal is to {goal} for compound {compound}.
Constraints:
- Budget: ${budget}
- Approved reagents only
- Complete within {time_h} hours
- Previous attempts: {previous}
Respond with structured JSON describing your protocol."""
import json, logging
from pathlib import Path
from typing import Dict, List, Any, Optional
from dataclasses import asdict
from functools import partial

MODEL_IDEATE   = "o4-mini-2025-04-16"  # o4-mini model for ideation - balances speed and quality

# Configure logging to help with tracking experiment progress and debugging
logging.basicConfig(level=logging.INFO, format="%(message)s")
logging.info(f"Run‑id {ctx.run_id}  Compound: {ctx.compound}")
logging.info(f"Logs will be stored in: {Path('logs') / ctx.run_id}")

def ideation(ctx: Context):
    logging.info("Starting ideation phase...")
    ideas = []
    for role, focus in ROLE_FOCUS.items():
        logging.info(f"Running ideation agent ${role}")
        sys = IDEATION_PROMPT.format(role=role, focus=focus, **ctx.prompt_vars())
        usr = f"Design a protocol to {ctx.goal} within ${ctx.budget}."
        idea = call_openai(ctx.client, MODEL_IDEATE, sys, usr, ctx)
        ideas.append(idea)
    log_json("ideation_done", ideas, ctx)
    return ideas
Run‑id 9835f69c  Compound: XYZ-13
Logs will be stored in: logs/9835f69c

Los agentes de ideación pueden utilizar herramientas externas como literature_search, chem_lookup (base de datos química), cost_estimator, outcome_db (resultado de experimentos anteriores) para fundamentar sus sugerencias en datos. Habilitar explícitamente y solicitar a los modelos que usen herramientas externas asegura que los planes generados sean factibles, conformes e informados por el conocimiento existente. El modelo decide cuándo y qué herramienta llamar según la tarea.

IDEATION_PROMPT += """\nUse the following tools as appropriate:
- Use the `list_available_chemicals` tool to get list of approved reagents.
- Use the `chem_lookup` tool to verify properties of reagents mentioned.
- Use the `cost_estimator` tool to calculate the approximate cost based on reagents and proposed steps.
- Check the `outcome_db` for relevant prior experiments with {compound}"""

ideas = ideation(ctx)
logging.info("Ideation complete!")
Starting ideation phase...
Running ideation agent $hypothesis_agent
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) List available chemicals
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Outcome DB: XYZ-13, yield, 5
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Cost estimator: [{'name': 'Palladium chloride', 'amount': 0.05, 'unit': 'g'}, {'name': 'Triphenylphosphine', 'amount': 0.1, 'unit': 'g'}, {'name': 'Potassium carbonate', 'amount': 1, 'unit': 'g'}, {'name': 'Dimethylformamide', 'amount': 50, 'unit': 'mL'}, {'name': 'Toluene', 'amount': 50, 'unit': 'mL'}, {'name': 'Sodium borohydride', 'amount': 0.1, 'unit': 'g'}, {'name': 'Triethylamine', 'amount': 0.5, 'unit': 'mL'}], ['round-bottom flask', 'magnetic stirrer', 'reflux condenser'], 36
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
Running ideation agent $protocol_agent
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Outcome DB: XYZ-13, yield, 5
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) List available chemicals
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Literature search: XYZ-13 synthesis palladium triphenylphosphine ligand yield improvement, None, 3
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Cost estimator: [{'name': 'Palladium acetate', 'amount': 0.05, 'unit': 'g'}, {'name': 'Triphenylphosphine', 'amount': 0.1, 'unit': 'g'}, {'name': 'Potassium carbonate', 'amount': 2, 'unit': 'g'}, {'name': 'Triethylamine', 'amount': 2, 'unit': 'mL'}, {'name': 'Dimethylformamide', 'amount': 100, 'unit': 'mL'}], ['Magnetic stirrer', 'Oil bath', 'Inert gas setup'], 48
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
Running ideation agent $resource_agent
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Outcome DB: XYZ-13, yield, 5
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) List available chemicals
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Cost estimator: [{'name': 'Palladium acetate', 'amount': 0.05, 'unit': 'g'}, {'name': 'Triphenylphosphine', 'amount': 0.1, 'unit': 'g'}, {'name': 'Potassium carbonate', 'amount': 1, 'unit': 'g'}, {'name': 'Dimethylformamide', 'amount': 5, 'unit': 'mL'}, {'name': 'Triethylamine', 'amount': 2, 'unit': 'mL'}], ['Round-bottom flask', 'Reflux condenser', 'Heating mantle', 'Magnetic stirrer'], 36
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Chemical lookup: Sodium borohydride, None
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
Ideation complete!

Estas herramientas se definen en agent_utils.py. Para los propósitos de esta solución, las llamadas a herramientas se simulan en tools.py. En un caso de uso real, estas herramientas llamarían a APIs reales.

2.3. Clasificación por torneo (o4-mini / o3):

Los protocolos generados se comparan por pares basándose en criterios como la efectividad esperada, la viabilidad, el costo y la novedad. En lugar de pedirle a un modelo que califique los protocolos de forma aislada, proporcionar dos protocolos a la vez y pedir una comparación directa con criterios específicos a menudo produce clasificaciones relativas más confiables.

Esta clasificación estilo Elo identifica a los candidatos más prometedores para una revisión más profunda.

TOURNAMENT_PROMPT = """
Protocol A: [details...]
Protocol B: [details...]

Compare Protocol A and Protocol B for synthesizing {compound} aimed at {goal}. Score them on:
1. Likelihood of achieving ≥ 15% yield increase.
2. Practical feasibility (reagents, time).
3. Estimated cost-efficiency (use tool if needed).
4. Scientific novelty/risk.

Return JSON {{\"winner\": \"A\"|\"B\", \"justification\": \"...\"}}."""

# This is a mock tourname implementation that only compares the first two protocols
# A real implementation would compare pairs in a tournament bracket style
def tournament(protocols: List[Dict[str, Any]], ctx: Context):
    logging.info("Starting tournament phase...")
    if len(protocols) == 1:
        return protocols[:1]
    a, b = protocols[0], protocols[1]
    sys = TOURNAMENT_PROMPT.format(**ctx.prompt_vars())
    usr = json.dumps({"A": a, "B": b}, indent=2)
    res = call_openai(ctx.client, MODEL_IDEATE, sys, usr, ctx)
    winner = a if res.get("winner", "A").upper() == "A" else b
    log_json("tournament", res, ctx)
    return [winner]

top_proto = tournament(ideas, ctx)[0]
logging.info("Tournament winner picked!")
Starting tournament phase...
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
Tournament winner picked!

En experimentos iniciales, descubrimos que pedir a los modelos que calificaran los protocolos en una escala de 1 a 10 conducía a resultados inconsistentes con la compresión de la puntuación. El enfoque de torneo resolvió esto al forzar juicios relativos que resultaron más confiables. Esto refleja el comportamiento de los expertos humanos: a los científicos a menudo les resulta más fácil comparar dos opciones directamente que asignar puntuaciones absolutas.

2.4. Crítica y síntesis profunda (o3):

Los protocolos mejor clasificados se pasan a o3 para una revisión rigurosa. o3 actúa como un científico sénior, evaluando la validez científica, la metodología, la seguridad, el cumplimiento del presupuesto y sugiriendo mejoras o sintetizando un protocolo final y refinado. También puede llamar a herramientas para su verificación.

# Deep critique phase using a more powerful model for rigorous review
CRITIQUE_PROMPT = """You are a senior researcher reviewing a proposed synthesis protocol 
for {compound} aiming for {goal}, budget ${budget} using approved reagents. Review the protocol below rigorously:
1. Identify scientific flaws or methodological weaknesses.
2. Assess safety risks and budget compliance (use `cost_estimator` tool if needed).
3. Check for consistency with prior `outcome_db` results if relevant.
4. Suggest concrete improvements or rewrite sections if necessary.
5. Provide a final go/no-go recommendation.

Return JSON {{\"revised_protocol\": ..., \"critique\": \"...\", \"recommendation\": \"go|no-go\"}}.

Protocol to Review:
[Protocol details...]
"""

MODEL_CRITIQUE = "o3-2025-04-16"  # o3 model for deep critique

def critique(protocol: Dict[str, Any], ctx: Context):
    logging.info("Starting critique phase...")
    sys = CRITIQUE_PROMPT.format(**ctx.prompt_vars())
    usr = json.dumps(protocol, indent=2)
    crit = call_openai(ctx.client, MODEL_CRITIQUE, sys, usr, ctx)
    log_json("critique", crit, ctx)
    return crit.get("revised_protocol", protocol)

critiqued = critique(top_proto, ctx)
logging.info("Deep critique completed!")
Starting critique phase...
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Cost estimator: [{'name': 'Palladium chloride', 'amount': 0.0045, 'unit': 'g'}, {'name': 'Triphenylphosphine', 'amount': 0.013, 'unit': 'g'}, {'name': 'Sodium borohydride', 'amount': 0.0038, 'unit': 'g'}, {'name': 'Potassium carbonate', 'amount': 0.14, 'unit': 'g'}, {'name': 'Triethylamine', 'amount': 0.07, 'unit': 'mL'}, {'name': 'Dimethylformamide', 'amount': 2, 'unit': 'mL'}, {'name': 'Toluene', 'amount': 5, 'unit': 'mL'}], ['100 mL round-bottom flask', 'magnetic stirrer', 'reflux condenser', 'inert gas line'], 24
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Outcome DB: XYZ-13, None, 5
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
Deep critique completed!

Deliberadamente separamos la ideación de la crítica utilizando diferentes modelos y personas. Que el mismo modelo genere y critique su propio trabajo a menudo lleva a la autojustificación en lugar de una evaluación objetiva. El modelo o3, actuando como un "científico sénior", identificó consistentemente debilidades metodológicas que o4-mini pasó por alto durante la ideación.

2.5. (Opcional) Verificación de seguridad:

Un modelo especializado, como gpt-4.1-mini, puede realizar una verificación final para preocupaciones de seguridad específicas (por ejemplo, combinaciones de reactivos peligrosos).

# Optional safety check using a targeted model
SAFETY_PROMPT = """You are a lab‑safety specialist. 
Identify hazards, unsafe conditions, or compliance issues in this protocol for {compound}. 
Use `chem_lookup` tool if needed. Return JSON assessment."""

MODEL_SAFETY   = "gpt-4.1-mini-2025-04-14"  # gpt-4.1-mini model for safety checks - optimized for instruction following

def safety(protocol: Dict[str, Any], ctx: Context):
    logging.info("Starting safety assessment...")
    sys = SAFETY_PROMPT.format(**ctx.prompt_vars())
    usr = json.dumps(protocol, indent=2)
    assessment = call_openai(ctx.client, MODEL_SAFETY, sys, usr, ctx)
    log_json("safety", assessment, ctx)
    return {"protocol": protocol, "safety": assessment}

secured = safety(critiqued, ctx)
logging.info("Safety check completed!")
Starting safety assessment...
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
(Tool) Chemical lookup: Palladium chloride, None
(Tool) Chemical lookup: Triphenylphosphine, None
(Tool) Chemical lookup: Sodium borohydride, None
(Tool) Chemical lookup: Potassium carbonate, None
(Tool) Chemical lookup: Dimethylformamide, None
(Tool) Chemical lookup: Toluene, None
HTTP Request: POST https://api.openai.com/v1/chat/completions "HTTP/1.1 200 OK"
Safety check completed!
Lección del curso «OpenAI Cookbook» de OpenAI, publicado con licencia MIT. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a OpenAI. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios