Lección 224 · 30 min · Gratis

Primeros pasos con la API de Responses de OpenAI en Amazon Bedrock (parte 2 de 2)

4.3 Manejar múltiples llamadas a herramientas

Las llamadas a herramientas paralelas permiten que el modelo solicite más de una búsqueda independiente en un solo turno. Esta celda permite dos búsquedas de estado de pedido, ejecuta cada llamada a función local y envía ambas salidas antes de pedirle al modelo que compare los problemas de envío activos. Inspecciona los ID de pedido devueltos y la respuesta final para confirmar que los datos de la aplicación, no la memoria del modelo, fundamentan la respuesta.

from __future__ import annotations
parallel_input = [{"role": "user", "content": "Use get_order_status for ORDER-8831 and ORDER-2044, then summarize whether Maya has one shipping problem or multiple active shipping problems in two labeled plain-text lines. Do not use leading hyphens or bold text."}]
parallel_request = {
    "model": MODEL_ID,
    "input": parallel_input,
    "tools": function_tools,
    "tool_choice": "auto",
    "parallel_tool_calls": True,
    "max_output_tokens": 1024,
    "store": False,
}

print_request_shape(parallel_request)
try:
    parallel_plan = create_response(**parallel_request)
    parallel_calls = [item for item in response_items(parallel_plan) if item.get("type") == "function_call"]
    parallel_outputs = [dispatch_tool_call(call) for call in parallel_calls]
    parallel_order_ids = [json.loads(call["arguments"]).get("order_id") for call in parallel_calls if call.get("name") == "get_order_status"]
    expected_order_ids = ["ORDER-8831", "ORDER-2044"]
    missing_order_ids = [order_id for order_id in expected_order_ids if order_id not in builtins.set(parallel_order_ids)]

    if not missing_order_ids:
        parallel_final = create_response(
            model=MODEL_ID,
            input=parallel_input + response_items(parallel_plan) + parallel_outputs,
            tools=function_tools,
            max_output_tokens=1024,
            store=False,
        )
        parallel_answer = output_text(parallel_final).strip()
        record_check("Parallel tool calls", "pass", {"tool_call_count": len(parallel_calls), "order_ids": parallel_order_ids})
        record_response("Parallel order lookup answer", "text", parallel_answer)
        print_labeled_json("Result: tool calls", {"tool_call_count": len(parallel_calls), "order_ids": parallel_order_ids})
        print_labeled_text("Result: final model answer", parallel_answer)
        print_response_summary(parallel_final)
        print_key_takeaway('Parallel tool calls let the model request multiple lookups, while the application still controls execution.')
    else:
        fallback_orders = [get_order_status(order_id) for order_id in expected_order_ids]
        fallback_prompt = (
            "The model did not request every expected order lookup. Use these application lookup results "
            "to answer in two labeled plain-text lines without leading hyphens or bold text: " + json.dumps(fallback_orders)
        )
        parallel_final = create_response(
            model=MODEL_ID,
            input=parallel_input + [{"role": "user", "content": fallback_prompt}],
            max_output_tokens=1024,
            store=False,
        )
        parallel_answer = output_text(parallel_final).strip()
        record_check("Parallel tool calls", "warn", {"returned_order_ids": parallel_order_ids, "missing_order_ids": missing_order_ids})
        record_response("Parallel order lookup fallback answer", "text", parallel_answer)
        print_labeled_json("Result: returned tool calls", {"tool_call_count": len(parallel_calls), "order_ids": parallel_order_ids})
        print_labeled_json("Result: local tool outputs", fallback_orders)
        print_labeled_text("Result: final model answer", parallel_answer)
        print_response_summary(parallel_final)
        print_key_takeaway('Local lookup outputs keep the parallel-tool pattern understandable if not every call is returned.')
except Exception as exc:
    handle_example_error("Parallel tool calls", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field value
model openai.gpt-5.4
max_output_tokens 1024
store False
parallel_tool_calls True
tools get_order_status, get_customer_profile
tool_choice auto
input 1 item(s): user: Use get_order_status for ORDER-8831 and ORDER-2044, then summarize whether Maya has one shipping problem or multiple active shipping problems in two labeled plain-text lines. Do not use leading hyphens or bold text.
<IPython.core.display.HTML object>
Resultado: llamadas a herramientas devueltas
{ "tool_call_count": 1, "order_ids": [ "ORDER-8831" ] }
<IPython.core.display.HTML object>
Resultado: salidas de herramientas locales
[ { "order_id": "ORDER-8831", "customer_id": "CUST-1042", "item": "standing desk replacement", "status": "delayed", "carrier_scan": "No movement for 36 hours at Denver sort center", "promised_delivery": "2026-06-01", "recommended_policy": "If delay exceeds 48 hours, offer expedited replacement or 15% concession with agent approval." }, { "order_id": "ORDER-2044", "customer_id": "CUST-1042", "item": "ergonomic chair", "status": "delivered", "carrier_scan": "Delivered yesterday at front desk", "promised_delivery": "2026-05-29", "recommended_policy": "Confirm delivery details before opening a replacement request." } ]
<IPython.core.display.HTML object>
Resultado: respuesta final del modelo
Order statuses: ORDER-8831 is delayed and ORDER-2044 was delivered yesterday. Shipping problems: Maya has one active shipping problem, because only ORDER-8831 is currently delayed while ORDER-2044 is already delivered.
<IPython.core.display.HTML object>
Resumen de la respuesta
<IPython.core.display.HTML object>
field value
id resp_ovmlijo2lf2nmc7udlofbbw7n2xmlxs6mxdpftlp6iamfcjoeerq
model openai.gpt-5.4
status completed
output_item_types ['message']
input_tokens 404
cached_input_tokens 0
output_tokens 50
total_tokens 454
reasoning_output_tokens 0
service_tier default
<IPython.core.display.HTML object>
Conclusión clave: Las salidas de búsqueda locales hacen que el patrón de herramientas paralelas sea comprensible, incluso si no se devuelve cada llamada.

4.4 Usar una herramienta de texto personalizada

Las herramientas personalizadas pasan texto de formato libre a la lógica propiedad de la aplicación en lugar de requerir un objeto de argumento JSON estructurado. Esta celda define un normalizador de notas de soporte, solicita una llamada a una herramienta personalizada e incluye un fallback local si el endpoint devuelve texto ordinario en lugar de una llamada personalizada. Inspecciona los tipos de elementos de salida y la nota normalizada.

from __future__ import annotations
custom_tools = [
    {
        "type": "custom",
        "name": "normalize_support_note",
        "description": "Normalize a freeform support note written by an agent. Input is plain text.",
        "format": {"type": "text"},
    }
]


def normalize_support_note_text(note: str) -> str:
    fields = [part.strip().upper() for part in note.split("|")]
    labels = ["ORDER_ID", "CUSTOMER_ID", "ISSUE", "CUSTOMER_REQUEST", "POLICY_OPTION"]
    return "\n".join(
        f"{label}: {value}"
        for label, value in zip(labels, fields)
        if value
    )


support_note = "order-8831 | cust-1042 | replacement delayed | customer wants supervisor | offer expedited replacement or 15% concession"
custom_input = [{
    "role": "user",
    "content": (
        "Call normalize_support_note with this exact note. Do not answer directly; "
        f"send the note to the custom tool: {support_note}"
    ),
}]
custom_request = {
    "model": MODEL_ID,
    "input": custom_input,
    "tools": custom_tools,
    "tool_choice": {"type": "custom", "name": "normalize_support_note"},
    "max_output_tokens": 1024,
    "store": False,
}

print_request_shape(custom_request)
print_labeled_text("Result: local fallback normalization", normalize_support_note_text(support_note))
try:
    custom_plan = create_response(**custom_request)
    returned_item_types = [item.get("type") for item in response_items(custom_plan)]
    try:
        custom_call = first_output_item(custom_plan, "custom_tool_call")
        if custom_call is None:
            raise LookupError("No custom_tool_call item returned.")
        tool_input = custom_call.get("input", "").strip()
        normalized_note = normalize_support_note_text(tool_input)
        record_check("Custom tools", "pass", {"output_item_types": returned_item_types, "normalized_note": normalized_note})
        record_response("Normalized support note", "text", normalized_note)
        print_labeled_json("Result: returned output item types", returned_item_types)
        print_labeled_text("Result: custom tool input", tool_input)
        print_labeled_text("Result: application-owned normalized output", normalized_note)
    except LookupError:
        fallback_text = output_text(custom_plan).strip() or "No text content was returned."
        normalized_note = normalize_support_note_text(support_note)
        record_check("Custom tools", "warn", {
            "expected": "custom_tool_call item named normalize_support_note",
            "actual_output_item_types": returned_item_types,
            "meaning": "The model response did not include a custom-tool invocation, so the application fallback normalization is shown for teaching.",
        })
        record_response("Custom tool text fallback", "text", fallback_text)
        record_response("Application-owned normalization fallback", "text", normalized_note)
        print_labeled_json("Result: returned output item types", returned_item_types or ["no typed output items returned"])
        print_labeled_text("Result: model text response", fallback_text)
        print_labeled_text("Result: application-owned normalization", normalized_note)
    print_response_summary(custom_plan)
    print_key_takeaway('Custom tools are useful when the application owns a freeform parsing or execution step.')
except Exception as exc:
    handle_example_error("Custom tools", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field value
model openai.gpt-5.4
max_output_tokens 1024
store False
tools normalize_support_note
tool_choice {'type': 'custom', 'name': 'normalize_support_note'}
input 1 item(s): user: Call normalize_support_note with this exact note. Do not answer directly; send the note to the custom tool: order-8831 | cust-1042 | replacement delayed | customer wants supervisor | offer expedited replacement or 15% co...
<IPython.core.display.HTML object>
Resultado: normalización de fallback local
ORDER_ID: ORDER-8831 CUSTOMER_ID: CUST-1042 ISSUE: REPLACEMENT DELAYED CUSTOMER_REQUEST: CUSTOMER WANTS SUPERVISOR POLICY_OPTION: OFFER EXPEDITED REPLACEMENT OR 15% CONCESSION
<IPython.core.display.HTML object>
Resultado: tipos de elementos de salida devueltos
[ "custom_tool_call" ]
<IPython.core.display.HTML object>
Resultado: entrada de herramienta personalizada
order-8831 | cust-1042 | replacement delayed | customer wants supervisor | offer expedited replacement or 15% concession
<IPython.core.display.HTML object>
Resultado: salida normalizada propiedad de la aplicación
ORDER_ID: ORDER-8831 CUSTOMER_ID: CUST-1042 ISSUE: REPLACEMENT DELAYED CUSTOMER_REQUEST: CUSTOMER WANTS SUPERVISOR POLICY_OPTION: OFFER EXPEDITED REPLACEMENT OR 15% CONCESSION
<IPython.core.display.HTML object>
Resumen de la respuesta
<IPython.core.display.HTML object>
field value
id resp_gwyauif44dnxpxrcssrxj4bh57tmgg3zwr67hfcklfswshbbscoa
model openai.gpt-5.4
status completed
output_item_types ['custom_tool_call']
input_tokens 674
cached_input_tokens 0
output_tokens 37
total_tokens 711
reasoning_output_tokens 0
service_tier default
<IPython.core.display.HTML object>
Conclusión clave: Las herramientas personalizadas son útiles cuando la aplicación posee un paso de análisis o ejecución de formato libre.

5. Enviar entrada de archivo directa

La entrada de archivo directa es independiente de las herramientas gestionadas por la aplicación. Se puede incluir un archivo en la solicitud de Responses actual como un elemento input_file junto con instrucciones de texto, lo cual es útil cuando el modelo debe leer el archivo para este turno sin configurar un índice de recuperación.

5.1 Adjuntar un PDF como input_file

Esta celda genera una pequeña transcripción PDF en memoria, la adjunta como datos de archivo base64 y solicita campos JSON exactos del documento. Inspecciona la vista previa del PDF, los campos esperados, la respuesta analizada y el resumen de uso.

from __future__ import annotations
def make_simple_pdf(lines: list[str]) -> bytes:
    def pdf_escape(text: str) -> str:
        return text.replace("\\", "\\\\").replace("(", "\\(").replace(")", "\\)")

    stream_lines = ["BT", "/F1 11 Tf", "72 740 Td", "15 TL"]
    for idx, line in enumerate(lines):
        if idx:
            stream_lines.append("T*")
        stream_lines.append(f"({pdf_escape(line)}) Tj")
    stream_lines.append("ET")
    stream = "\n".join(stream_lines).encode("latin-1", "replace")

    objects = [
        b"<< /Type /Catalog /Pages 2 0 R >>",
        b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
        b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] /Resources << /Font << /F1 4 0 R >> >> /Contents 5 0 R >>",
        b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica >>",
        b"<< /Length " + builtins.str(len(stream)).encode("ascii") + b" >>\nstream\n" + stream + b"\nendstream",
    ]

    pdf = b"%PDF-1.4\n"
    offsets = [0]
    for idx, obj in enumerate(objects, start=1):
        offsets.append(len(pdf))
        pdf += f"{idx} 0 obj\n".encode("ascii") + obj + b"\nendobj\n"
    xref_offset = len(pdf)
    pdf += f"xref\n0 {len(objects) + 1}\n0000000000 65535 f \n".encode("ascii")
    for offset in offsets[1:]:
        pdf += f"{offset:010d} 00000 n \n".encode("ascii")
    pdf += f"trailer\n<< /Size {len(objects) + 1} /Root 1 0 R >>\nstartxref\n{xref_offset}\n%%EOF\n".encode("ascii")
    return pdf


file_lines = [
    "BrightCart support transcript",
    "Ticket: TICKET-7429",
    "Customer: Maya Chen",
    "Order: ORDER-8831",
    "Product: Standing desk replacement",
    "Issue: Replacement for a damaged item is delayed and carrier scan has not moved",
    "Customer request: Supervisor callback and refund options",
    "Policy options: expedited replacement or 15% concession with agent approval after 48-hour delay",
]
file_text = "\n".join(file_lines)
pdf_data = base64.b64encode(make_simple_pdf(file_lines)).decode("utf-8")

expected_direct_file_fields = {
    "ticket_id": "TICKET-7429",
    "customer": "Maya Chen",
    "order_id": "ORDER-8831",
    "product": "Standing desk replacement",
}

direct_file_request = {
    "model": MODEL_ID,
    "input": [
        {
            "role": "user",
            "content": [
                {
                    "type": "input_file",
                    "filename": "brightcart-support-transcript.pdf",
                    "file_data": f"data:application/pdf;base64,{pdf_data}",
                },
                {
                    "type": "input_text",
                    "text": (
                        "Read the attached PDF support transcript and return JSON with keys "
                        "ticket_id, customer, order_id, product, issue, requested_resolution, and policy_options. "
                        "Use exact values from the file. Do not return null for fields that are present in the file."
                    ),
                },
            ],
        }
    ],
    "text": {"format": {"type": "json_object"}},
    "max_output_tokens": 1024,
    "store": False,
}

print_labeled_text("Result: PDF transcript preview", file_text)
print_request_shape(direct_file_request)
print_labeled_json("Result: expected fields", expected_direct_file_fields)
try:
    direct_file_response = create_response(**direct_file_request)
    raw_direct_file_output = output_text(direct_file_response).strip()
    try:
        direct_file_payload = json.loads(raw_direct_file_output)
        missing_or_empty = [
            key for key, expected in expected_direct_file_fields.items()
            if builtins.str(direct_file_payload.get(key, "")).strip().lower() != expected.lower()
        ]
        null_fields = [key for key, value in direct_file_payload.items() if value in {None, "", []}]
        if missing_or_empty or null_fields:
            record_check("Direct file inputs", "warn", {
                "message": "The request completed, but the model did not extract the expected values from the attached PDF.",
                "missing_or_unexpected_fields": missing_or_empty,
                "empty_fields": null_fields,
                "payload": direct_file_payload,
            })
            record_response("Support transcript extraction returned by model", "json", direct_file_payload)
            print_labeled_text("Result", "The request completed, but the model did not extract the expected values from the attached PDF.")
            print_labeled_json("Result: returned JSON", direct_file_payload)
        else:
            record_check("Direct file inputs", "pass", direct_file_payload)
            record_response("Support transcript extraction", "json", direct_file_payload)
            print_labeled_json("Result", direct_file_payload)
    except Exception as parse_exc:
        record_check("Direct file inputs", "warn", {
            "message": "The request completed, but the response was not valid JSON.",
            "text_sample": raw_direct_file_output[:600],
            "error": builtins.str(parse_exc),
        })
        record_response("Support transcript extraction text", "text", raw_direct_file_output[:1200])
        print_labeled_text("Result", raw_direct_file_output[:1200])
    print_response_summary(direct_file_response)
    print_key_takeaway('Direct file input is useful when the file should be read in the current request context.')
except Exception as exc:
    handle_example_error("Direct file inputs", exc)
<IPython.core.display.HTML object>
Resultado: vista previa de la transcripción PDF
BrightCart support transcript Ticket: TICKET-7429 Customer: Maya Chen Order: ORDER-8831 Product: Standing desk replacement Issue: Replacement for a damaged item is delayed and carrier scan has not moved Customer request: Supervisor callback and refund options Policy options: expedited replacement or 15% concession with agent approval after 48-hour delay
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field value
model openai.gpt-5.4
max_output_tokens 1024
store False
text format json_object
input 1 item(s): user: input_file: brightcart-support-transcript.pdf; input_text: Read the attached PDF support transcript and return JSON with keys ticket_id, customer, order_id, product, issue, reques...
<IPython.core.display.HTML object>
Resultado: campos esperados
{ "ticket_id": "TICKET-7429", "customer": "Maya Chen", "order_id": "ORDER-8831", "product": "Standing desk replacement" }
<IPython.core.display.HTML object>
Resultado
{"ticket_id":"TICKET-7429","customer":"Maya Chen","order_id":"ORDER-8831","product":"Standing desk replacement","issue":"Replacement for a damaged item is delayed and carrier scan has not moved","requested_resolution":"Supervisor callback and refund options","policy_options":"expedited replacement or 15% concession with agent approval after 48-hour delay"}
<IPython.core.display.HTML object>
Resumen de la respuesta
<IPython.core.display.HTML object>
field value
id resp_colsvndmpjd6qczpemdjscmsbjefmgl5vh6i7alqt52jflentfna
model openai.gpt-5.4
status completed
output_item_types ['message']
input_tokens 713
cached_input_tokens 0
output_tokens 82
total_tokens 795
reasoning_output_tokens 0
service_tier default
<IPython.core.display.HTML object>
Conclusión clave: La entrada de archivo directa es útil cuando el archivo debe leerse en el contexto de la solicitud actual.

6. Gestionar el estado de la conversación

El estado de la conversación determina cómo los turnos de seguimiento reciben el contexto anterior. La API de Responses admite la continuación almacenada con previous_response_id, y las aplicaciones también pueden gestionar el estado por sí mismas reenviando el historial de entrada relevante. Esta sección compara ambos patrones y luego muestra el contexto de razonamiento cifrado donde se admite.

6.1 Continuar con previous_response_id

Usa previous_response_id para continuar desde una respuesta almacenada sin reenviar el prompt completo anterior. La primera solicitud almacena los detalles del caso BrightCart; la segunda solicitud pasa solo la nueva instrucción de seguimiento más el ID de respuesta anterior. Inspecciona si el seguimiento conserva el pedido, el cliente, el problema y la siguiente acción.

from __future__ import annotations
promised_delivery = (date.today() + timedelta(days=2)).isoformat()
stateful_seed_input = (
    f"Customer Maya Chen opened ticket TICKET-4812 about order ORDER-8831. "
    "The item is a standing desk replacement for a damaged delivery. "
    f"The promised delivery date is {promised_delivery}, but the carrier scan has not moved in 36 hours. "
    "Customer sentiment is frustrated because this is the second attempt. "
    "Support policy says to offer expedited replacement or a 15% concession if the delay exceeds 48 hours. "
    "Escalation owner is Tier 2 Returns."
)
stateful_followup_input = "Return five labeled lines: ticket ID, order ID, customer name, issue, and next best action."
stateful_request_shape = {
    "model": MODEL_ID,
    "input": stateful_followup_input,
    "previous_response_id": "<response-id-from-prior-stored-turn>",
    "max_output_tokens": 1024,
    "store": False,
}

print_request_shape(stateful_request_shape)
try:
    stateful_turn_1 = create_response(model=MODEL_ID, input=stateful_seed_input, max_output_tokens=1024, store=True)
    remember_stored_response(stateful_turn_1)
    stateful_turn_2 = create_response(model=MODEL_ID, input=stateful_followup_input, previous_response_id=stateful_turn_1.id, max_output_tokens=1024, store=False)
    text = output_text(stateful_turn_2).strip()
    require("order-8831" in text.lower() or "maya" in text.lower(), "Stateful continuation response missed expected support context.")
    record_check("Stateful continuation", "pass", stateful_turn_1.id)
    record_response("Stateful support handoff", "text", text)
    print_labeled_text("Result", text)
    print_response_summary(stateful_turn_2)
    print_key_takeaway('previous_response_id lets a follow-up use stored context without resending the full prior turn.')
except Exception as exc:
    handle_example_error("Stateful continuation", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field value
model openai.gpt-5.4
max_output_tokens 1024
store False
previous_response_id <response-id-from-prior-stored-turn>
input Return five labeled lines: ticket ID, order ID, customer name, issue, and next best action.
<IPython.core.display.HTML object>
Resultado
Ticket ID: TICKET-4812 Order ID: ORDER-8831 Customer Name: Maya Chen Issue: Replacement standing desk shipment for damaged delivery has had no carrier movement for 36 hours; customer is frustrated because this is the second attempt Next Best Action: Monitor until 48 hours without movement, then offer expedited replacement or 15% concession and escalate to Tier 2 Returns if needed
<IPython.core.display.HTML object>
Resumen de la respuesta
<IPython.core.display.HTML object>
field value
id resp_gkgqadc2gd24747lmhy5waftt5tga7eibtv67k77lndjgbuioo6q
model openai.gpt-5.4
status completed
output_item_types ['message']
input_tokens 715
cached_input_tokens 0
output_tokens 86
total_tokens 801
reasoning_output_tokens 0
service_tier default
<IPython.core.display.HTML object>
Conclusión clave: previous_response_id permite que un seguimiento use el contexto almacenado sin reenviar el turno anterior completo.

6.2 Reconstruir el contexto sin estado

La continuación sin estado significa que la aplicación envía el historial relevante en cada solicitud. Esto es adecuado cuando tu producto ya posee el almacenamiento de conversaciones, la política de retención o los requisitos de auditoría. Esta celda envía un historial de chat corto más una nueva instrucción de traspaso e inspecciona el resumen y el uso de tokens.

from __future__ import annotations
stateless_history = [
    {"role": "user", "content": "Support chat TICKET-3920: Customer Jordan Lee says ORDER-7718 arrived with a cracked monitor stand."},
    {"role": "assistant", "content": "Captured damaged-item issue for ORDER-7718 and asked for preferred resolution."},
    {"role": "user", "content": "Jordan wants a replacement shipped this week and asks whether the damaged item must be returned first."},
]
stateless_payload = {
    "model": MODEL_ID,
    "input": stateless_history + [{"role": "user", "content": "Summarize this support chat for the next agent in five labeled plain-text lines. Do not use leading hyphens or bold text."}],
    "max_output_tokens": 1024,
    "store": False,
}

print_request_shape(stateless_payload)
try:
    stateless_response = create_response(**stateless_payload)
    stateless_text = output_text(stateless_response).strip()
    require(stateless_text, "Stateless continuation response did not return text.")
    record_check("Stateless continuation", "pass", summarize_response(stateless_response))
    record_response("Stateless support handoff", "text", stateless_text)
    print_labeled_text("Result", stateless_text)
    print_response_summary(stateless_response)
    print_key_takeaway('Stateless continuation sends the relevant history with each request when the application owns conversation storage.')
except Exception as exc:
    handle_example_error("Stateless continuation", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field value
model openai.gpt-5.4
max_output_tokens 1024
store False
input 4 item(s): user: Support chat TICKET-3920: Customer Jordan Lee says ORDER-7718 arrived with a cracked monitor stand.; assistant: Captured damaged-item issue for ORDER-7718 and asked for preferred resolution.; user: Jordan wants a replacement shipped this week and asks whether the damaged item must be returned first.; user: Summarize this support chat for the next agent in five labeled plain-text lines. Do not use leading hyphens or bold text.
<IPython.core.display.HTML object>
Resultado
Customer: Jordan Lee reported ORDER-7718 arrived with a cracked monitor stand. Issue: Damaged item; monitor stand is cracked on arrival. Requested Resolution: Customer wants a replacement shipped this week. Open Question: Jordan asked whether the damaged item must be returned before replacement is sent. Status: Damage claim captured and awaiting next-agent confirmation on replacement timing and return requirement.
<IPython.core.display.HTML object>
Resumen de la respuesta
<IPython.core.display.HTML object>
field value
id resp_mezt6yqizyswuvujvudnonr34b73ndyyu2qsfncgtjyppzie5vva
model openai.gpt-5.4
status completed
output_item_types ['message']
input_tokens 255
cached_input_tokens 0
output_tokens 78
total_tokens 333
reasoning_output_tokens 0
service_tier default
<IPython.core.display.HTML object>
Conclusión clave: La continuación sin estado envía el historial relevante con cada solicitud cuando la aplicación posee el almacenamiento de conversaciones.

6.3 Llevar el contexto de razonamiento cifrado

Los modelos con capacidad de razonamiento pueden devolver elementos de razonamiento y contenido de razonamiento cifrado cuando se les solicita. Esta celda solicita metadatos de razonamiento cifrados, lleva elementos de respuesta anteriores a una solicitud de seguimiento e inspecciona si se devolvió contenido cifrado. El texto de razonamiento oculto no se expone; la aplicación solo lleva el contexto opaco hacia adelante cuando es compatible.

Documentos oficiales: Modelos de razonamiento describe los modelos de razonamiento y el esfuerzo de razonamiento en los flujos de trabajo de respuestas.

from __future__ import annotations
encrypted_history = [
    {"role": "user", "content": "For a customer-support assistant handling names, order IDs, and refund context, compare stateful and stateless continuation in two sentences."}
]
encrypted_turn_payload = {
    "model": MODEL_ID,
    "input": encrypted_history,
    "reasoning": {"effort": "medium"},
    "include": ["reasoning.encrypted_content"],
    "max_output_tokens": 1024,
    "store": False,
}

print_request_shape(encrypted_turn_payload)
try:
    encrypted_turn_1 = create_response(**encrypted_turn_payload)
    encrypted_turn_2 = create_response(
        model=MODEL_ID,
        input=encrypted_history + response_items(encrypted_turn_1) + [
            {"role": "user", "content": "Based on the prior reasoning context, recommend one approach for a regulated support workflow in two labeled plain-text lines. Do not use leading hyphens or bold text."}
        ],
        max_output_tokens=1024,
        store=False,
    )
    reasoning_items = [item for item in response_items(encrypted_turn_1) if item.get("type") == "reasoning"]
    has_encrypted_content = any(item.get("encrypted_content") for item in reasoning_items)
    record_check("Encrypted reasoning", "pass", {"encrypted_content_returned": has_encrypted_content, "reasoning_item_count": len(reasoning_items)})
    encrypted_answer = output_text(encrypted_turn_2).strip()
    record_response("State strategy recommendation", "text", encrypted_answer)
    print_labeled_json("Result: reasoning metadata", {
        "returned_item_types": [item.get("type") for item in response_items(encrypted_turn_1)],
        "encrypted_reasoning_content_returned": has_encrypted_content,
    })
    print_labeled_text("Result: follow-up answer", encrypted_answer)
    print_response_summary(encrypted_turn_2)
    print_key_takeaway('Encrypted reasoning content can be carried forward where supported without exposing hidden reasoning text.')
except Exception as exc:
    handle_example_error("Encrypted reasoning", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field value
model openai.gpt-5.4
max_output_tokens 1024
store False
reasoning {'effort': 'medium'}
include ['reasoning.encrypted_content']
input 1 item(s): user: Para un asistente de atención al cliente que maneja nombres, ID de pedidos y contexto de reembolsos, compara la continuación con estado y sin estado en dos oraciones.
<IPython.core.display.HTML object>
Resultado: metadatos de razonamiento
{ "returned_item_types": [ "reasoning", "message" ], "encrypted_reasoning_content_returned": true }
<IPython.core.display.HTML object>
Resultado: respuesta de seguimiento
Recomendación: Continuación sin estado Razón: En un flujo de trabajo de soporte regulado, requerir que los nombres, ID de pedidos y el contexto de reembolso se proporcionen explícitamente en cada turno mejora la controlabilidad, la auditabilidad y la minimización de datos, reduciendo el riesgo de retención no intencionada o fuga entre sesiones.
<IPython.core.display.HTML object>
Resumen de la respuesta
<IPython.core.display.HTML object>
field value
id resp_qmjgoymsxqf32ht3apisbvscrv4d5t5x2tnxeftkxwpyl4eppjka
model openai.gpt-5.4
status completed
output_item_types ['message']
input_tokens 295
cached_input_tokens 0
output_tokens 56
total_tokens 351
reasoning_output_tokens 0
service_tier default
<IPython.core.display.HTML object>
Conclusión clave: El contenido de razonamiento cifrado se puede llevar adelante donde sea compatible sin exponer el texto de razonamiento oculto.

7. Usar el almacenamiento en caché de prompts

El almacenamiento en caché de prompts mejora la latencia y el costo cuando las solicitudes comparten un prefijo estático exacto.

7.1 Comparar dos solicitudes con clave de caché

Esta celda coloca texto de política estable de BrightCart al principio de la entrada, envía la misma solicitud dos veces con un prompt_cache_key y compara los metadatos de los tokens. Inspecciona cached_input_tokens en la segunda respuesta cuando el endpoint devuelve los detalles de la caché.

Nota: PROMPT_CACHE_RETENTION se selecciona del MODEL_ID activo. Usa 24h para openai.gpt-5.5 y modelos posteriores, y in_memory para openai.gpt-5.4.

from __future__ import annotations
base_support_policy = [
    "BrightCart support policy:",
    "1. Be empathetic, concise, and specific about the customer's order.",
    "2. Do not promise refunds, credits, or delivery dates unless the policy context supports it.",
    "3. For damaged-item replacements, check replacement status before offering concessions.",
    "4. If a replacement delay exceeds 48 hours, offer expedited replacement or a 15% concession subject to agent approval.",
]
policy_reference_paragraph = (
    "Expanded cacheable policy context: BrightCart agents should identify the customer, order ID, replacement status, "
    "carrier scan age, promised delivery window, item category, prior concessions, and supervisor approval needs before "
    "drafting a customer-facing answer. The assistant should preserve a calm tone, avoid unsupported promises, separate "
    "confirmed facts from assumptions, recommend one clear next action, and document why any escalation, expedited "
    "replacement, or concession is appropriate. Repeated policy context like this is intentionally stable across many "
    "requests so prompt caching can reuse the prefix when the same cache key is supplied."
)
expanded_policy_context = "\n".join(
    f"Policy reference paragraph {idx + 1}: {policy_reference_paragraph}"
    for idx in range(32)
)
stable_support_policy = "\n".join(base_support_policy + [expanded_policy_context])
cache_input = [
    {"role": "system", "content": stable_support_policy},
    {"role": "user", "content": "Draft a two-sentence agent reply for Maya Chen about delayed replacement order ORDER-8831."},
]
estimated_cache_input_words = len(json.dumps(cache_input).split())
require(estimated_cache_input_words > 2048, f"Prompt-cache input should be over 2048 words; found {estimated_cache_input_words}.")
cache_payload = {
    "model": MODEL_ID,
    "input": cache_input,
    "prompt_cache_key": "brightcart-support-policy-v1",
    "prompt_cache_retention": PROMPT_CACHE_RETENTION,
    "max_output_tokens": 1024,
    "store": False,
}

print_request_shape(cache_payload)
print_labeled_json("Prompt-cache input size", {"estimated_input_words": estimated_cache_input_words, "target_minimum_tokens": 2048})
try:
    cache_response_1 = create_response(**cache_payload)
    cache_response_2 = create_response(**cache_payload)
    cache_summary_1 = summarize_response(cache_response_1)
    cache_summary_2 = summarize_response(cache_response_2)
    cache_comparison = pd.DataFrame([
        {
            "request": "first",
            "input_tokens": cache_summary_1.get("input_tokens"),
            "cached_input_tokens": cache_summary_1.get("cached_input_tokens"),
            "output_tokens": cache_summary_1.get("output_tokens"),
            "total_tokens": cache_summary_1.get("total_tokens"),
        },
        {
            "request": "second",
            "input_tokens": cache_summary_2.get("input_tokens"),
            "cached_input_tokens": cache_summary_2.get("cached_input_tokens"),
            "output_tokens": cache_summary_2.get("output_tokens"),
            "total_tokens": cache_summary_2.get("total_tokens"),
        },
    ])
    record_check("Prompt caching", "pass" if cache_summary_2.get("cached_input_tokens") is not None else "warn", {"first": cache_summary_1, "second": cache_summary_2})
    cache_reply = output_text(cache_response_2).strip()
    record_response("Prompt-cache token comparison", "table", cache_comparison)
    record_response("Cached support-policy reply", "text", cache_reply)

    print_labeled_text("Result", cache_reply)
    print_labeled_json("First request summary", cache_summary_1)
    print_labeled_json("Second request summary", cache_summary_2)
    print_label("Response summary")
    display_wrapped_table(cache_comparison, max_col_width_px=260)
    print_key_takeaway('cached_input_tokens is the metadata field to inspect for prompt-cache reuse.')
except Exception as exc:
    handle_example_error("Prompt caching", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field value
model openai.gpt-5.4
max_output_tokens 1024
store False
prompt_cache_key brightcart-support-policy-v1
prompt_cache_retention in_memory
input 2 item(s): system: Política de soporte de BrightCart: 1. Sé empático, conciso y específico sobre el pedido del cliente. 2. No prometas reembolsos, créditos o fechas de entrega a menos que el contexto de la política lo respalde. 3. Para reemplazos de artículos dañados...; user: Redacta una respuesta de agente de dos oraciones para Maya Chen sobre el pedido de reemplazo retrasado ORDER-8831.
<IPython.core.display.HTML object>
Tamaño de entrada de la caché de prompts
{ "estimated_input_words": 3016, "target_minimum_tokens": 2048 }
<IPython.core.display.HTML object>
Resultado
Hola Maya, lamento que tu pedido de reemplazo ORDER-8831 esté retrasado. Estoy verificando el estado más reciente del reemplazo y del transportista para poder confirmar el mejor siguiente paso para ti sin hacerte esperar más de lo necesario.
<IPython.core.display.HTML object>
Resumen de la primera solicitud
{ "id": "resp_3w6r6ipbqa5z2max35awv3i23i5sjmfa33zzw3vhpgqxcvchhkgq", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": [ "message" ], "input_tokens": 3970, "output_tokens": 66, "total_tokens": 4036, "cached_input_tokens": 0, "reasoning_output_tokens": 0, "service_tier": "default" }
<IPython.core.display.HTML object>
Resumen de la segunda solicitud
{ "id": "resp_zzjeqttoswdjdwwl56xpolvly23w4h2n5dtsdoddgxqwfbkn7npq", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": [ "message" ], "input_tokens": 3970, "output_tokens": 48, "total_tokens": 4018, "cached_input_tokens": 0, "reasoning_output_tokens": 0, "service_tier": "default" }
<IPython.core.display.HTML object>
Resumen de la respuesta
<IPython.core.display.HTML object>
request input_tokens cached_input_tokens output_tokens total_tokens
first 3970 0 66 4036
second 3970 0 48 4018
<IPython.core.display.HTML object>
Conclusión clave: cached_input_tokens es el campo de metadatos a inspeccionar para la reutilización de la caché de prompts.

8. Ejecutar trabajo en segundo plano

El modo en segundo plano inicia una respuesta de forma asíncrona y permite que la aplicación consulte el estado final.

8.1 Enviar y consultar una respuesta en segundo plano

Esta celda envía background=true, almacena el ID de la respuesta, consulta mientras el estado está en cola o en progreso, y luego imprime el resumen final del gerente. Inspecciona el historial de estado, el estado final, el ID de la respuesta y el resumen de tokens.

from __future__ import annotations
backlog = """
Same-day BrightCart support backlog:
1. 18 delayed-order contacts, mostly from the West Coast distribution lane.
2. 7 damaged-item replacement contacts; 3 mention replacement delays.
3. 5 return-window exception requests after holiday promotions.
""".strip()
background_payload = {
    "model": MODEL_ID,
    "input": f"Return exactly three labeled plain-text lines for a support-manager summary: theme, risk, next action. Keep each line under 12 words. Do not use leading hyphens or bold text.\n\n{backlog}",
    "background": True,
    "max_output_tokens": 1024,
    "store": True,
}

print_request_shape(background_payload)
try:
    background_response = create_response(**background_payload)
    remember_stored_response(background_response)
    status_history = [getattr(background_response, "status", None)]
    for _ in range(15):
        if getattr(background_response, "status", None) not in {"queued", "in_progress"}:
            break
        time.sleep(2)
        background_response = retrieve_response(background_response.id)
        status_history.append(getattr(background_response, "status", None))
    background_summary = summarize_response(background_response)
    manager_summary = output_text(background_response).strip()
    require(manager_summary, "Background response did not return text.")
    status = "pass" if background_summary.get("status") in {None, "completed"} else "warn"
    record_check("Background mode", status, {"status_history": status_history, "id": getattr(background_response, "id", None), "final_status": background_summary.get("status")})
    record_response("Background manager summary", "text", manager_summary)
    print_labeled_json("Result: status history", status_history)
    print_labeled_text("Result: manager summary", manager_summary)
    print_response_summary(background_summary)
    print_key_takeaway('Background mode starts work asynchronously and lets the application poll by response ID.')
except Exception as exc:
    handle_example_error("Background mode", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field value
model openai.gpt-5.4
max_output_tokens 1024
store True
background True
input Devuelve exactamente tres líneas de texto plano etiquetadas para un resumen del gerente de soporte: tema, riesgo, siguiente acción. Mantén cada línea por debajo de 12 palabras. No uses guiones iniciales ni texto en negrita. Retraso de soporte de BrightCart del mismo día: 1. 18 contactos de pedidos retrasados, la mayoría de la We...
<IPython.core.display.HTML object>
Resultado: historial de estado
[ "in_progress", "completed" ]
<IPython.core.display.HTML object>
Resultado: resumen del gerente
tema: Los retrasos en el envío dominan, especialmente en la ruta de distribución de la Costa Oeste. riesgo: Aumento de la insatisfacción por retrasos, reemplazos y excepciones de devolución. siguiente acción: Escalar los problemas de la ruta de la Costa Oeste y revisar la política de devoluciones de vacaciones.
<IPython.core.display.HTML object>
Resumen de la respuesta
<IPython.core.display.HTML object>
field value
id resp_lmmtsvgk3ntolh5ci5vxccmsa6uxcgrsq7v54jpz7oewmociyesa
model openai.gpt-5.4
status completed
output_item_types ['message']
input_tokens 246
cached_input_tokens 0
output_tokens 45
total_tokens 291
reasoning_output_tokens 0
service_tier default
<IPython.core.display.HTML object>
Conclusión clave: El modo en segundo plano inicia el trabajo de forma asíncrona y permite que la aplicación consulte por ID de respuesta.

9. Compactar el contexto de larga duración

La compactación reduce el estado de una conversación larga a hechos duraderos, preguntas abiertas, restricciones y próximas acciones. Esta celda documenta el patrón de compactación del lado de la aplicación como un pequeño objeto JSON para que el concepto sea claro sin agregar otra ruta de función en vivo. Inspecciona qué hechos se mantienen y qué detalles se omiten antes del siguiente turno.

from __future__ import annotations
compaction_note = {
    "feature": "Compaction",
    "how_to_apply": "Summarize older support turns into durable facts, open questions, policy constraints, and next actions before continuing the workflow.",
    "brightcart_example": {
        "durable_facts": ["Customer Maya Chen", "ORDER-8831", "replacement delayed", "carrier scan stale"],
        "policy_constraints": ["Do not promise refund without eligibility", "Offer expedited replacement or 15% concession after 48-hour delay with approval"],
        "next_action": "Check latest carrier scan and supervisor callback status.",
    },
}
record_check("Compaction", "documented", compaction_note)
record_response("Compacted support context", "json", compaction_note)
print_json(compaction_note)
<IPython.core.display.HTML object>
JSON
{ "feature": "Compaction", "how_to_apply": "Resume los turnos de soporte anteriores en hechos duraderos, preguntas abiertas, restricciones de política y próximas acciones antes de continuar el flujo de trabajo.", "brightcart_example": { "durable_facts": [ "Cliente Maya Chen", "ORDER-8831", "reemplazo retrasado", "escaneo del transportista obsoleto" ], "policy_constraints": [ "No prometas reembolso sin elegibilidad", "Ofrece reemplazo expedito o concesión del 15% después de un retraso de 48 horas con aprobación" ], "next_action": "Verifica el último escaneo del transportista y el estado de la llamada del supervisor." } }

10. Ejecutar verificaciones operativas rápidas (Smoke Checks)

Las verificaciones operativas rápidas son comprobaciones de configuración ligeras, no una prueba de carga o una medición a nivel de servicio. Esta celda envía tres solicitudes cortas, mide el tiempo transcurrido local, resume la tasa de éxito y el uso de tokens, e infiere la región de la URL base de Bedrock configurada. Inspecciona la latencia, el estado de finalización, las salidas de muestra y los totales de tokens.

from __future__ import annotations
def infer_region_from_base_url(base_url: str) -> str | None:
    host = normalize_base_url(base_url).replace("https://", "").split("/")[0]
    for part in host.split("."):
        if part.count("-") >= 2 and any(char.isdigit() for char in part):
            return part
    return None


def percentile(values: list[float], pct: float) -> float | None:
    if not values:
        return None
    ordered = sorted(values)
    index = min(len(ordered) - 1, max(0, round((pct / 100) * (len(ordered) - 1))))
    return round(ordered[index], 3)


operations_features = ["Latency runtime example", "Throughput runtime example", "Reliability runtime example", "Region check"]
operations_payload = {"model": MODEL_ID, "input": "Reply with one short customer-support sentence.", "service_tier": "auto", "max_output_tokens": 1024, "store": False}

print_request_shape(operations_payload)
if not RUN_RESPONSIVENESS_CHECK:
    record_check("Endpoint responsiveness", "skipped", "BEDROCK_RESPONSIVENESS_CHECK is disabled.")
    print_labeled_text("Result", "Responsiveness check disabled.")
else:
    prompts = [
        "Reply in one short sentence: apologize for a delayed replacement order.",
        "Reply with one metric name for support-assistant quality.",
        "Reply in one short sentence: hand off a return exception to a supervisor.",
    ]
    samples = []
    for idx, prompt in enumerate(prompts):
        started = time.perf_counter()
        try:
            response = create_response(model=MODEL_ID, input=prompt, service_tier="auto", max_output_tokens=1024, store=False)
            elapsed = time.perf_counter() - started
            summary = summarize_response(response)
            text = output_text(response).strip()
            samples.append({
                "ok": bool(text),
                "latency_seconds": round(elapsed, 3),
                "output_tokens": summary.get("output_tokens") or 0,
                "total_tokens": summary.get("total_tokens") or 0,
                "status": summary.get("status"),
                "sample_output": text[:140],
            })
        except Exception as exc:
            elapsed = time.perf_counter() - started
            samples.append({"ok": False, "latency_seconds": round(elapsed, 3), "error": describe_api_error(exc)})

    successes = [sample for sample in samples if sample["ok"]]
    completed = [sample for sample in successes if sample.get("status") in {None, "completed"}]
    latencies = [sample["latency_seconds"] for sample in successes]
    responsiveness_summary = {
        "region_hint": infer_region_from_base_url(BASE_URL),
        "base_url_host": normalize_base_url(BASE_URL).replace("https://", "").split("/")[0],
        "sample_count": len(samples),
        "success_rate": len(successes) / len(samples) if samples else 0,
        "completed_rate": len(completed) / len(samples) if samples else 0,
        "avg_latency_seconds": round(sum(latencies) / len(latencies), 3) if latencies else None,
        "p50_latency_seconds": percentile(latencies, 50),
        "p90_latency_seconds": percentile(latencies, 90),
        "total_output_tokens": sum(sample.get("output_tokens", 0) for sample in samples),
        "total_tokens": sum(sample.get("total_tokens", 0) for sample in samples),
    }
    status = "pass" if len(successes) == len(samples) and len(completed) == len(samples) else "warn"
    for feature in operations_features:
        record_check(feature, status, responsiveness_summary)
    record_response("Endpoint responsiveness summary", "json", {**responsiveness_summary, "samples": samples})
    print_labeled_json("Result", responsiveness_summary)
    print_label("Response summary")
    display_wrapped_table(pd.DataFrame(samples), max_col_width_px=360)
    print_key_takeaway('Responsiveness samples are setup checks, not a load test or service-level measurement.')
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field value
model openai.gpt-5.4
max_output_tokens 1024
store False
service_tier auto
input Responde con una oración corta de atención al cliente.
<IPython.core.display.HTML object>
Resultado
{ "region_hint": "us-west-2", "base_url_host": "bedrock-mantle.us-west-2.api.aws", "sample_count": 3, "success_rate": 1.0, "completed_rate": 1.0, "avg_latency_seconds": 0.362, "p50_latency_seconds": 0.377, "p90_latency_seconds": 0.4, "total_output_tokens": 34, "total_tokens": 544 }
<IPython.core.display.HTML object>
Resumen de la respuesta
<IPython.core.display.HTML object>
ok latency_seconds output_tokens total_tokens status sample_output
True 0.400 14 184 completed Lamentamos el retraso con su pedido de reemplazo.
True 0.310 6 174 completed Tasa de resolución
True 0.377 14 186 completed Estoy escalando la excepción de devolución a un supervisor.
<IPython.core.display.HTML object>
Conclusión clave: Las muestras de capacidad de respuesta son verificaciones de configuración, no una prueba de carga o una medición a nivel de servicio.

11. Limpia y revisa los resultados

Las respuestas almacenadas creadas por los ejemplos de ciclo de vida, continuación con estado y segundo plano se rastrean en STORED_RESPONSE_IDS. Esta celda final intenta eliminar las respuestas almacenadas cuando la limpieza está habilitada, luego imprime el resumen de la ejecución y la galería de respuestas de ejemplo. Primero, inspecciona las advertencias; estas suelen identificar diferencias en la configuración del endpoint, la disponibilidad del modelo o el soporte de características.

from __future__ import annotations

cleanup_rows = []
tracked_ids = list(dict.fromkeys(STORED_RESPONSE_IDS))

if not tracked_ids:
    cleanup_rows.append({"response_id": "none", "status": "no stored responses tracked", "detail": ""})
elif not CLEAN_UP_STORED_RESPONSES:
    for stored_id in tracked_ids:
        cleanup_rows.append({"response_id": stored_id, "status": "skipped", "detail": "BEDROCK_CLEANUP_STORED_RESPONSES is disabled"})
    record_check("Stored response cleanup", "skipped", cleanup_rows)
else:
    for stored_id in tracked_ids:
        try:
            delete_result = delete_response(stored_id)
            cleanup_rows.append({"response_id": stored_id, "status": "deleted", "detail": compact_text(to_dict(delete_result), 240)})
        except Exception as exc:
            cleanup_rows.append({"response_id": stored_id, "status": "warn", "detail": describe_api_error(exc)})
    cleanup_status = "pass" if all(row["status"] == "deleted" for row in cleanup_rows) else "warn"
    record_check("Stored response cleanup", cleanup_status, cleanup_rows)

print_label("Stored response cleanup")
display_wrapped_table(pd.DataFrame(cleanup_rows), max_col_width_px=520)

summary_df = pd.DataFrame(RESULTS_SUMMARY)
print_label("Run summary")
display_wrapped_table(summary_df, max_col_width_px=620)

print_label("Example responses")
print_response_gallery()
<IPython.core.display.HTML object>
Limpieza de respuestas almacenadas
<IPython.core.display.HTML object>
response_id status detail
resp_cvhvh7y5ghwrpa35snvk4bzgcgthgxp4tgwkllmf5mrhs7dikfia warn {'exception_class': 'AuthenticationError', 'status_code': 401, 'retryable': False, 'request_id': 'req_gkni5zyr7lkjkz2vfiwvkev2qgxs76crcwz5whhjrdkma7up3yta', 'message': 'Error code: 401 - {'error': {'code': 'invalid_api_key', 'message': 'The security token included in the request is invalid.', 'param': None, 'type': 'permission_denied_error'}}'}
resp_vjrtvnakcgxjhnq5b7cj7rtowtdh7chkkf3aqwbdkynhjqiklp3a warn {'exception_class': 'AuthenticationError', 'status_code': 401, 'retryable': False, 'request_id': 'req_vvwhsmp2rkrzwbkajqdpdod2o4j5xo2vxelzxbcp2fjan2ybwc2a', 'message': 'Error code: 401 - {'error': {'code': 'invalid_api_key', 'message': 'The security token included in the request is invalid.', 'param': None, 'type': 'permission_denied_error'}}'}
resp_lmmtsvgk3ntolh5ci5vxccmsa6uxcgrsq7v54jpz7oewmociyesa warn {'exception_class': 'AuthenticationError', 'status_code': 401, 'retryable': False, 'request_id': 'req_btgbpxspnokm3wzfndybnfuv3kjudxt3r2ihlsmru7ziisn52goa', 'message': 'Error code: 401 - {'error': {'code': 'invalid_api_key', 'message': 'The security token included in the request is invalid.', 'param': None, 'type': 'permission_denied_error'}}'}
<IPython.core.display.HTML object>
Resumen de la ejecución
<IPython.core.display.HTML object>
name status detail
Endpoint shape pass https://bedrock-mantle.us-west-2.api.aws/openai/v1/responses
Model selection pass Using configured model; model-list metadata is not required for requests.
Error handling pass {"normalized_fields": ["exception_class", "status_code", "retryable", "request_id", "message"], "retryable_status_codes": [408, 409, 429, 500, 502, 503, 504], "notes": "call_with_retries(...) uses this taxonomy for transient retry handling."}
Text generation pass resp_naythl6fvzhoctlsdogd4vpr673q5ibagqqpiujbast3sy6viroa
Text generation pass {"id": "resp_nmvqefzghd5hi67uwy4wfvhwnnzild3lslqsxqor3cat63kmucoq", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["reasoning", "message"], "input_tokens": 177, "output_tokens": 129, "total_tokens": 306, "cached_input_tokens": 0, "reasoning_output_tokens": 18, "service_tier": "default"}
Reasoning effort pass {"id": "resp_nmvqefzghd5hi67uwy4wfvhwnnzild3lslqsxqor3cat63kmucoq", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["reasoning", "message"], "input_tokens": 177, "output_tokens": 129, "total_tokens": 306, "cached_input_tokens": 0, "reasoning_output_tokens": 18, "service_tier": "default"}
Responses lifecycle pass resp_cvhvh7y5ghwrpa35snvk4bzgcgthgxp4tgwkllmf5mrhs7dikfia
Response schema pass {"id": "resp_cvhvh7y5ghwrpa35snvk4bzgcgthgxp4tgwkllmf5mrhs7dikfia", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["message"], "input_tokens": 198, "output_tokens": 109, "total_tokens": 307, "cached_input_tokens": 0, "reasoning_output_tokens": 0, "service_tier": "default"}
Usage metadata pass {"id": "resp_cvhvh7y5ghwrpa35snvk4bzgcgthgxp4tgwkllmf5mrhs7dikfia", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["message"], "input_tokens": 198, "output_tokens": 109, "total_tokens": 307, "cached_input_tokens": 0, "reasoning_output_tokens": 0, "service_tier": "default"}
Prompt caching pass {"id": "resp_q4akwbeynfwfwnt5i4tdwkpcgsffdu4lqvng7lnwaob53opswvwq", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["reasoning", "message"], "input_tokens": 183, "output_tokens": 91, "total_tokens": 274, "cached_input_tokens": 0, "reasoning_output_tokens": 34, "service_tier": "default"}
Service tier pass {"id": "resp_q4akwbeynfwfwnt5i4tdwkpcgsffdu4lqvng7lnwaob53opswvwq", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["reasoning", "message"], "input_tokens": 183, "output_tokens": 91, "total_tokens": 274, "cached_input_tokens": 0, "reasoning_output_tokens": 34, "service_tier": "default"}
Reasoning effort pass {"id": "resp_q4akwbeynfwfwnt5i4tdwkpcgsffdu4lqvng7lnwaob53opswvwq", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["reasoning", "message"], "input_tokens": 183, "output_tokens": 91, "total_tokens": 274, "cached_input_tokens": 0, "reasoning_output_tokens": 34, "service_tier": "default"}
Structured Outputs pass {"ticket_id": "TICKET-7429", "category": "delivery_delay", "priority": "urgent", "customer_sentiment": "frustrated and time-sensitive", "summary": "Customer Maya Chen reports that ORDER-8831 is a replacement shipment for a previously damaged standing desk. The replacement is now 2 days late, carrier tracking has not updated, and she needs the desk delivered before Monday. She is requesting a supervisor callback and wants to know refund options if the replacement cannot arrive in time.", "require...
JSON mode pass {"customer_name": "Maya Chen", "order_id": "ORDER-8831", "issue_summary": "Customer is asking about a delayed replacement order. The carrier tracking scan is stale and has not updated.", "next_step": "Handoff to support to investigate the carrier delay, verify shipment status, and provide Maya Chen with an update or resolution.", "metrics_to_watch": ["tracking_scan_recency", "carrier_exception_status", "replacement_order_delivery_eta", "customer_follow_up_time"]}
Verbosity pass {"compact_chars": 182, "detailed_chars": 332}
Function calling pass {"tool_choice_used": "required", "arguments": {"order_id": "ORDER-8831"}}
Parallel tool calls warn {"returned_order_ids": ["ORDER-8831"], "missing_order_ids": ["ORDER-2044"]}
Custom tools pass {"output_item_types": ["custom_tool_call"], "normalized_note": "ORDER_ID: ORDER-8831\nCUSTOMER_ID: CUST-1042\nISSUE: REPLACEMENT DELAYED\nCUSTOMER_REQUEST: CUSTOMER WANTS SUPERVISOR\nPOLICY_OPTION: OFFER EXPEDITED REPLACEMENT OR 15% CONCESSION"}
Direct file inputs warn {"message": "The request completed, but the response was not valid JSON.", "text_sample": "{\"ticket_id\":\"TICKET-7429\",\"customer\":\"Maya Chen\",\"order_id\":\"ORDER-8831\",\"product\":\"Standing desk replacement\",\"issue\":\"Replacement for a damaged item is delayed and carrier scan has not moved\",\"requested_resolution\":\"Supervisor callback and refund options\",\"policy_options\":\"expedited replacement or 15% concession with agent approval after 48-hour delay\"}", "error": "unhashable...
Stateful continuation pass resp_vjrtvnakcgxjhnq5b7cj7rtowtdh7chkkf3aqwbdkynhjqiklp3a
Stateless continuation pass {"id": "resp_mezt6yqizyswuvujvudnonr34b73ndyyu2qsfncgtjyppzie5vva", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["message"], "input_tokens": 255, "output_tokens": 78, "total_tokens": 333, "cached_input_tokens": 0, "reasoning_output_tokens": 0, "service_tier": "default"}
Encrypted reasoning pass {"encrypted_content_returned": true, "reasoning_item_count": 1}
Prompt caching pass {"first": {"id": "resp_3w6r6ipbqa5z2max35awv3i23i5sjmfa33zzw3vhpgqxcvchhkgq", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["message"], "input_tokens": 3970, "output_tokens": 66, "total_tokens": 4036, "cached_input_tokens": 0, "reasoning_output_tokens": 0, "service_tier": "default"}, "second": {"id": "resp_zzjeqttoswdjdwwl56xpolvly23w4h2n5dtsdoddgxqwfbkn7npq", "model": "openai.gpt-5.4", "status": "completed", "output_item_types": ["message"], "input_tokens": 3970, "outp...
Background mode pass {"status_history": ["in_progress", "completed"], "id": "resp_lmmtsvgk3ntolh5ci5vxccmsa6uxcgrsq7v54jpz7oewmociyesa", "final_status": "completed"}
Compaction documented {"feature": "Compaction", "how_to_apply": "Summarize older support turns into durable facts, open questions, policy constraints, and next actions before continuing the workflow.", "brightcart_example": {"durable_facts": ["Customer Maya Chen", "ORDER-8831", "replacement delayed", "carrier scan stale"], "policy_constraints": ["Do not promise refund without eligibility", "Offer expedited replacement or 15% concession after 48-hour delay with approval"], "next_action": "Check latest carrier scan and...
Latency runtime example pass {"region_hint": "us-west-2", "base_url_host": "bedrock-mantle.us-west-2.api.aws", "sample_count": 3, "success_rate": 1.0, "completed_rate": 1.0, "avg_latency_seconds": 0.362, "p50_latency_seconds": 0.377, "p90_latency_seconds": 0.4, "total_output_tokens": 34, "total_tokens": 544}
Throughput runtime example pass {"region_hint": "us-west-2", "base_url_host": "bedrock-mantle.us-west-2.api.aws", "sample_count": 3, "success_rate": 1.0, "completed_rate": 1.0, "avg_latency_seconds": 0.362, "p50_latency_seconds": 0.377, "p90_latency_seconds": 0.4, "total_output_tokens": 34, "total_tokens": 544}
Reliability runtime example pass {"region_hint": "us-west-2", "base_url_host": "bedrock-mantle.us-west-2.api.aws", "sample_count": 3, "success_rate": 1.0, "completed_rate": 1.0, "avg_latency_seconds": 0.362, "p50_latency_seconds": 0.377, "p90_latency_seconds": 0.4, "total_output_tokens": 34, "total_tokens": 544}
Region check pass {"region_hint": "us-west-2", "base_url_host": "bedrock-mantle.us-west-2.api.aws", "sample_count": 3, "success_rate": 1.0, "completed_rate": 1.0, "avg_latency_seconds": 0.362, "p50_latency_seconds": 0.377, "p90_latency_seconds": 0.4, "total_output_tokens": 34, "total_tokens": 544}
Stored response cleanup warn [{"response_id": "resp_cvhvh7y5ghwrpa35snvk4bzgcgthgxp4tgwkllmf5mrhs7dikfia", "status": "warn", "detail": {"exception_class": "AuthenticationError", "status_code": 401, "retryable": false, "request_id": "req_gkni5zyr7lkjkz2vfiwvkev2qgxs76crcwz5whhjrdkma7up3yta", "message": "Error code: 401 - {'error': {'code': 'invalid_api_key', 'message': 'The security token included in the request is invalid.', 'param': None, 'type': 'permission_denied_error'}}"}}, {"response_id": "resp_vjrtvnakcgxjhnq5b7cj7rt...
<IPython.core.display.HTML object>
Respuestas de ejemplo
<IPython.core.display.HTML object>
example response_type response
Endpoint verification text ok
First raw HTTPS request text Empathy: I’m sorry, Maya — your replacement order ORDER-8831 is delayed because the carrier reported a temporary transit hold at the regional sorting facility.\nAction: We’re monitoring the shipment closely and will send you an updated delivery estimate within 24 hours; if there’s no movement by then, we’ll review the next replacement or refund options with you.
SDK text generation text Use the Responses API to build a BrightCart support assistant that can answer customer questions, summarize policies, and guide users through common workflows like order tracking, refunds, and account updates. Ground the assistant in BrightCart documentation and connect it to relevant backend tools or APIs so it can retrieve live order data, check account status, and provide accurate, context-aware support responses. Design the experience around clear system instructions, structured tool calling, and conversation state management so the assistant stays on-brand, reliable, and safe when handling customer issues.
Create and retrieve response text goal: Help support agents explain delayed replacement orders, set expectations, and suggest next steps.\ndata needed: Order ID, replacement order status, shipment/tracking events, delay reason, estimated ship/delivery date, customer contact history, inventory/backorder status, and applicable refund or reship policy.\nhuman-review rule: Escalate to a human if the delay exceeds policy thresholds, tracking is inconsistent or missing, the order appears lost, the customer is high-risk or highly upset, or any refund/reship exception is requested.
Service tier and prompt cache request text Latency benefit: Prompt caching lets the BrightCart support assistant reuse previously processed context, reducing response time for repeated or similar requests.\nConsistency benefit: Prompt caching helps the BrightCart support assistant return more uniform answers by reusing the same established prompt context across interactions.
Structured ticket triage json {\n  "ticket_id": "TICKET-7429",\n  "category": "delivery_delay",\n  "priority": "urgent",\n  "customer_sentiment": "frustrated and time-sensitive",\n  "summary": "Customer Maya Chen reports that ORDER-8831 is a replacement shipment for a previously damaged standing desk. The replacement is now 2 days late, carrier tracking has not updated, and she needs the desk delivered before Monday. She is requesting a supervisor callback and wants to know refund options if the replacement cannot arrive in time.",\n  "required_actions": [\n    "Review ORDER-8831 shipment status and confirm last carrier scan/update.",\n    "Contact carrier or open a trace/escalation for stalled tracking.",\n    "Check expedited reshipment or alternative fulfillment options to meet the before-Monday deadline.",\n    "Arrange supervisor callback per customer request.",\n    "Review and communicate refund options, including refund\n...
JSON support handoff json {\n  "customer_name": "Maya Chen",\n  "order_id": "ORDER-8831",\n  "issue_summary": "Customer is asking about a delayed replacement order. The carrier tracking scan is stale and has not updated.",\n  "next_step": "Handoff to support to investigate the carrier delay, verify shipment status, and provide Maya Chen with an update or resolution.",\n  "metrics_to_watch": [\n    "tracking_scan_recency",\n    "carrier_exception_status",\n    "replacement_order_delivery_eta",\n    "customer_follow_up_time"\n  ]\n}
Compact policy guidance text BrightCart’s delayed-replacement policy lets customers keep using the original item until the replacement arrives, then return the defective product within the allowed return window.
Detailed policy guidance text 1. BrightCart sends replacements after customers return the original item and warehouse receipt is confirmed.\n2. This delay prevents duplicate shipments, verifies eligibility, and reduces fraud or inventory errors.\n3. Agents should explain timelines clearly, offer return instructions, and reassure customers once receipt is logged.
Order-status tool answer text Status: ORDER-8831 is delayed; carrier shows no movement for 36 hours at the Denver sort center, with promised delivery on 2026-06-01.\nNext best action: Monitor until the 48-hour threshold; if no movement then, contact the customer and offer either an expedited replacement or a 15% concession with agent approval.
Parallel order lookup fallback answer text Order statuses: ORDER-8831 is delayed and ORDER-2044 was delivered yesterday.\nShipping problems: Maya has one active shipping problem, because only ORDER-8831 is currently delayed while ORDER-2044 is already delivered.
Normalized support note text ORDER_ID: ORDER-8831\nCUSTOMER_ID: CUST-1042\nISSUE: REPLACEMENT DELAYED\nCUSTOMER_REQUEST: CUSTOMER WANTS SUPERVISOR\nPOLICY_OPTION: OFFER EXPEDITED REPLACEMENT OR 15% CONCESSION
Support transcript extraction text text {"ticket_id":"TICKET-7429","customer":"Maya Chen","order_id":"ORDER-8831","product":"Standing desk replacement","issue":"Replacement for a damaged item is delayed and carrier scan has not moved","requested_resolution":"Supervisor callback and refund options","policy_options":"expedited replacement or 15% concession with agent approval after 48-hour delay"}
Stateful support handoff text Ticket ID: TICKET-4812\nOrder ID: ORDER-8831\nCustomer Name: Maya Chen\nIssue: Replacement standing desk shipment for damaged delivery has had no carrier movement for 36 hours; customer is frustrated because this is the second attempt\nNext Best Action: Monitor until 48 hours without movement, then offer expedited replacement or 15% concession and escalate to Tier 2 Returns if needed
Stateless support handoff text Customer: Jordan Lee reported ORDER-7718 arrived with a cracked monitor stand.\nIssue: Damaged item; monitor stand is cracked on arrival.\nRequested Resolution: Customer wants a replacement shipped this week.\nOpen Question: Jordan asked whether the damaged item must be returned before replacement is sent.\nStatus: Damage claim captured and awaiting next-agent confirmation on replacement timing and return requirement.
State strategy recommendation text Recommendation: Stateless continuation\nReason: In a regulated support workflow, requiring names, order IDs, and refund context to be explicitly provided each turn improves controllability, auditability, and data-minimization, reducing the risk of unintended retention or cross-session leakage.
Prompt-cache token comparison table [\n  {\n    "request":"first",\n    "input_tokens":3970,\n    "cached_input_tokens":0,\n    "output_tokens":66,\n    "total_tokens":4036\n  },\n  {\n    "request":"second",\n    "input_tokens":3970,\n    "cached_input_tokens":0,\n    "output_tokens":48,\n    "total_tokens":4018\n  }\n]
Cached support-policy reply text Hi Maya, I’m sorry your replacement order ORDER-8831 is delayed. I’m checking the latest replacement and carrier status now so I can confirm the best next step for you without making you wait longer than necessary.
Background manager summary text theme: Shipping delays dominate, especially West Coast distribution lane.\nrisk: Rising dissatisfaction from delays, replacements, and return exceptions.\nnext action: Escalate West Coast lane issues and review holiday return policy.
Compacted support context json {\n  "feature": "Compaction",\n  "how_to_apply": "Summarize older support turns into durable facts, open questions, policy constraints, and next actions before continuing the workflow.",\n  "brightcart_example": {\n    "durable_facts": [\n      "Customer Maya Chen",\n      "ORDER-8831",\n      "replacement delayed",\n      "carrier scan stale"\n    ],\n    "policy_constraints": [\n      "Do not promise refund without eligibility",\n      "Offer expedited replacement or 15% concession after 48-hour delay with approval"\n    ],\n    "next_action": "Check latest carrier scan and supervisor callback status."\n  }\n}
Endpoint responsiveness summary json {\n  "region_hint": "us-west-2",\n  "base_url_host": "bedrock-mantle.us-west-2.api.aws",\n  "sample_count": 3,\n  "success_rate": 1.0,\n  "completed_rate": 1.0,\n  "avg_latency_seconds": 0.362,\n  "p50_latency_seconds": 0.377,\n  "p90_latency_seconds": 0.4,\n  "total_output_tokens": 34,\n  "total_tokens": 544,\n  "samples": [\n    {\n      "ok": true,\n      "latency_seconds": 0.4,\n      "output_tokens": 14,\n      "total_tokens": 184,\n      "status": "completed",\n      "sample_output": "We apologize for the delay with your replacement order."\n    },\n    {\n      "ok": true,\n      "latency_seconds": 0.31,\n      "output_tokens": 6,\n      "total_tokens": 174,\n      "status": "completed",\n      "sample_output": "Resolution Rate"\n    },\n    {\n      "ok": true,\n      "latency_seconds": 0.377,\n      "output_tokens": 14,\n      "total_tokens": 186,\n      "status": "completed",\n      "sample_output": "I\u2019m e\n...
example response_type  \
0                   Endpoint verification          text   
1                 First raw HTTPS request          text   
2                     SDK text generation          text   
3            Create and retrieve response          text   
4   Service tier and prompt cache request          text   
5                Structured ticket triage          json   
6                    JSON support handoff          json   
7                 Compact policy guidance          text   
8                Detailed policy guidance          text   
9                Order-status tool answer          text   
10  Parallel order lookup fallback answer          text   
11                Normalized support note          text   
12     Support transcript extraction text          text   
13               Stateful support handoff          text   
14              Stateless support handoff          text   
15          State strategy recommendation          text   
16          Prompt-cache token comparison         table   
17            Cached support-policy reply          text   
18             Background manager summary          text   
19              Compacted support context          json   
20        Endpoint responsiveness summary          json   

                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                       response  
0                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              
… (salida recortada)
example response_type response
0 Endpoint verification text ok
1 First raw HTTPS request text Empathy: I’m sorry, Maya — your replacement order ORDER-8831 is delayed because the carrier reported a temporary transit hold at the regional sorting facility.\nAction: We’re monitoring the shipment closely and will send you an updated delivery estimate within 24 hours; if there’s no movement by then, we’ll review the next replacement or refund options with you.
2 SDK text generation text Use the Responses API to build a BrightCart support assistant that can answer customer questions, summarize policies, and guide users through common workflows like order tracking, refunds, and account updates. Ground the assistant in BrightCart documentation and connect it to relevant backend tools or APIs so it can retrieve live order data, check account status, and provide accurate, context-aware support responses. Design the experience around clear system instructions, structured tool calling, and conversation state management so the assistant stays on-brand, reliable, and safe when handling customer issues.
3 Create and retrieve response text goal: Help support agents explain delayed replacement orders, set expectations, and suggest next steps.\ndata needed: Order ID, replacement order status, shipment/tracking events, delay reason, estimated ship/delivery date, customer contact history, inventory/backorder status, and applicable refund or reship policy.\nhuman-review rule: Escalate to a human if the delay exceeds policy thresholds, tracking is inconsistent or missing, the order appears lost, the customer is high-risk or highly upset, or any refund/reship exception is requested.
4 Service tier and prompt cache request text Latency benefit: Prompt caching lets the BrightCart support assistant reuse previously processed context, reducing response time for repeated or similar requests.\nConsistency benefit: Prompt caching helps the BrightCart support assistant return more uniform answers by reusing the same established prompt context across interactions.
5 Structured ticket triage json {\n  "ticket_id": "TICKET-7429",\n  "category": "delivery_delay",\n  "priority": "urgent",\n  "customer_sentiment": "frustrated and time-sensitive",\n  "summary": "Customer Maya Chen reports that ORDER-8831 is a replacement shipment for a previously damaged standing desk. The replacement is now 2 days late, carrier tracking has not updated, and she needs the desk delivered before Monday. She is requesting a supervisor callback and wants to know refund options if the replacement cannot arrive in time.",\n  "required_actions": [\n    "Review ORDER-8831 shipment status and confirm last carrier scan/update.",\n    "Contact carrier or open a trace/escalation for stalled tracking.",\n    "Check expedited reshipment or alternative fulfillment options to meet the before-Monday deadline.",\n    "Arrange supervisor callback per customer request.",\n    "Review and communicate refund options, including refund\n...
6 JSON support handoff json {\n  "customer_name": "Maya Chen",\n  "order_id": "ORDER-8831",\n  "issue_summary": "Customer is asking about a delayed replacement order. The carrier tracking scan is stale and has not updated.",\n  "next_step": "Handoff to support to investigate the carrier delay, verify shipment status, and provide Maya Chen with an update or resolution.",\n  "metrics_to_watch": [\n    "tracking_scan_recency",\n    "carrier_exception_status",\n    "replacement_order_delivery_eta",\n    "customer_follow_up_time"\n  ]\n}
7 Compact policy guidance text BrightCart’s delayed-replacement policy lets customers keep using the original item until the replacement arrives, then return the defective product within the allowed return window.
8 Detailed policy guidance text 1. BrightCart sends replacements after customers return the original item and warehouse receipt is confirmed.\n2. This delay prevents duplicate shipments, verifies eligibility, and reduces fraud or inventory errors.\n3. Agents should explain timelines clearly, offer return instructions, and reassure customers once receipt is logged.
9 Order-status tool answer text Status: ORDER-8831 is delayed; carrier shows no movement for 36 hours at the Denver sort center, with promised delivery on 2026-06-01.\nNext best action: Monitor until the 48-hour threshold; if no movement then, contact the customer and offer either an expedited replacement or a 15% concession with agent approval.
10 Parallel order lookup fallback answer text Order statuses: ORDER-8831 is delayed and ORDER-2044 was delivered yesterday.\nShipping problems: Maya has one active shipping problem, because only ORDER-8831 is currently delayed while ORDER-2044 is already delivered.
11 Normalized support note text ORDER_ID: ORDER-8831\nCUSTOMER_ID: CUST-1042\nISSUE: REPLACEMENT DELAYED\nCUSTOMER_REQUEST: CUSTOMER WANTS SUPERVISOR\nPOLICY_OPTION: OFFER EXPEDITED REPLACEMENT OR 15% CONCESSION
12 Support transcript extraction text text {"ticket_id":"TICKET-7429","customer":"Maya Chen","order_id":"ORDER-8831","product":"Standing desk replacement","issue":"Replacement for a damaged item is delayed and carrier scan has not moved","requested_resolution":"Supervisor callback and refund options","policy_options":"expedited replacement or 15% concession with agent approval after 48-hour delay"}
13 Stateful support handoff text Ticket ID: TICKET-4812\nOrder ID: ORDER-8831\nCustomer Name: Maya Chen\nIssue: Replacement standing desk shipment for damaged delivery has had no carrier movement for 36 hours; customer is frustrated because this is the second attempt\nNext Best Action: Monitor until 48 hours without movement, then offer expedited replacement or 15% concession and escalate to Tier 2 Returns if needed
14 Stateless support handoff text Customer: Jordan Lee reported ORDER-7718 arrived with a cracked monitor stand.\nIssue: Damaged item; monitor stand is cracked on arrival.\nRequested Resolution: Customer wants a replacement shipped this week.\nOpen Question: Jordan asked whether the damaged item must be returned before replacement is sent.\nStatus: Damage claim captured and awaiting next-agent confirmation on replacement timing and return requirement.
15 State strategy recommendation text Recommendation: Stateless continuation\nReason: In a regulated support workflow, requiring names, order IDs, and refund context to be explicitly provided each turn improves controllability, auditability, and data-minimization, reducing the risk of unintended retention or cross-session leakage.
16 Prompt-cache token comparison table [\n  {\n    "request":"first",\n    "input_tokens":3970,\n    "cached_input_tokens":0,\n    "output_tokens":66,\n    "total_tokens":4036\n  },\n  {\n    "request":"second",\n    "input_tokens":3970,\n    "cached_input_tokens":0,\n    "output_tokens":48,\n    "total_tokens":4018\n  }\n]
17 Cached support-policy reply text Hi Maya, I’m sorry your replacement order ORDER-8831 is delayed. I’m checking the latest replacement and carrier status now so I can confirm the best next step for you without making you wait longer than necessary.
18 Background manager summary text theme: Shipping delays dominate, especially West Coast distribution lane.\nrisk: Rising dissatisfaction from delays, replacements, and return exceptions.\nnext action: Escalate West Coast lane issues and review holiday return policy.
19 Compacted support context json {\n  "feature": "Compaction",\n  "how_to_apply": "Summarize older support turns into durable facts, open questions, policy constraints, and next actions before continuing the workflow.",\n  "brightcart_example": {\n    "durable_facts": [\n      "Customer Maya Chen",\n      "ORDER-8831",\n      "replacement delayed",\n      "carrier scan stale"\n    ],\n    "policy_constraints": [\n      "Do not promise refund without eligibility",\n      "Offer expedited replacement or 15% concession after 48-hour delay with approval"\n    ],\n    "next_action": "Check latest carrier scan and supervisor callback status."\n  }\n}
20 Endpoint responsiveness summary json {\n  "region_hint": "us-west-2",\n  "base_url_host": "bedrock-mantle.us-west-2.api.aws",\n  "sample_count": 3,\n  "success_rate": 1.0,\n  "completed_rate": 1.0,\n  "avg_latency_seconds": 0.362,\n  "p50_latency_seconds": 0.377,\n  "p90_latency_seconds": 0.4,\n  "total_output_tokens": 34,\n  "total_tokens": 544,\n  "samples": [\n    {\n      "ok": true,\n      "latency_seconds": 0.4,\n      "output_tokens": 14,\n      "total_tokens": 184,\n      "status": "completed",\n      "sample_output": "We apologize for the delay with your replacement order."\n    },\n    {\n      "ok": true,\n      "latency_seconds": 0.31,\n      "output_tokens": 6,\n      "total_tokens": 174,\n      "status": "completed",\n      "sample_output": "Resolution Rate"\n    },\n    {\n      "ok": true,\n      "latency_seconds": 0.377,\n      "output_tokens": 14,\n      "total_tokens": 186,\n      "status": "completed",\n      "sample_output": "I\u2019m e\n...
Lección del curso «OpenAI Cookbook» de OpenAI, publicado con licencia MIT. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a OpenAI. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios
← AnteriorSiguiente: Un manual para la adopción segura y escalable de agentes de IA →