Primeros pasos con la API de Responses de OpenAI en Amazon Bedrock (parte 2 de 2)
4.3 Manejar múltiples llamadas a herramientas
Las llamadas a herramientas paralelas permiten que el modelo solicite más de una búsqueda independiente en un solo turno. Esta celda permite dos búsquedas de estado de pedido, ejecuta cada llamada a función local y envía ambas salidas antes de pedirle al modelo que compare los problemas de envío activos. Inspecciona los ID de pedido devueltos y la respuesta final para confirmar que los datos de la aplicación, no la memoria del modelo, fundamentan la respuesta.
from __future__ import annotations
parallel_input = [{"role": "user", "content": "Use get_order_status for ORDER-8831 and ORDER-2044, then summarize whether Maya has one shipping problem or multiple active shipping problems in two labeled plain-text lines. Do not use leading hyphens or bold text."}]
parallel_request = {
"model": MODEL_ID,
"input": parallel_input,
"tools": function_tools,
"tool_choice": "auto",
"parallel_tool_calls": True,
"max_output_tokens": 1024,
"store": False,
}
print_request_shape(parallel_request)
try:
parallel_plan = create_response(**parallel_request)
parallel_calls = [item for item in response_items(parallel_plan) if item.get("type") == "function_call"]
parallel_outputs = [dispatch_tool_call(call) for call in parallel_calls]
parallel_order_ids = [json.loads(call["arguments"]).get("order_id") for call in parallel_calls if call.get("name") == "get_order_status"]
expected_order_ids = ["ORDER-8831", "ORDER-2044"]
missing_order_ids = [order_id for order_id in expected_order_ids if order_id not in builtins.set(parallel_order_ids)]
if not missing_order_ids:
parallel_final = create_response(
model=MODEL_ID,
input=parallel_input + response_items(parallel_plan) + parallel_outputs,
tools=function_tools,
max_output_tokens=1024,
store=False,
)
parallel_answer = output_text(parallel_final).strip()
record_check("Parallel tool calls", "pass", {"tool_call_count": len(parallel_calls), "order_ids": parallel_order_ids})
record_response("Parallel order lookup answer", "text", parallel_answer)
print_labeled_json("Result: tool calls", {"tool_call_count": len(parallel_calls), "order_ids": parallel_order_ids})
print_labeled_text("Result: final model answer", parallel_answer)
print_response_summary(parallel_final)
print_key_takeaway('Parallel tool calls let the model request multiple lookups, while the application still controls execution.')
else:
fallback_orders = [get_order_status(order_id) for order_id in expected_order_ids]
fallback_prompt = (
"The model did not request every expected order lookup. Use these application lookup results "
"to answer in two labeled plain-text lines without leading hyphens or bold text: " + json.dumps(fallback_orders)
)
parallel_final = create_response(
model=MODEL_ID,
input=parallel_input + [{"role": "user", "content": fallback_prompt}],
max_output_tokens=1024,
store=False,
)
parallel_answer = output_text(parallel_final).strip()
record_check("Parallel tool calls", "warn", {"returned_order_ids": parallel_order_ids, "missing_order_ids": missing_order_ids})
record_response("Parallel order lookup fallback answer", "text", parallel_answer)
print_labeled_json("Result: returned tool calls", {"tool_call_count": len(parallel_calls), "order_ids": parallel_order_ids})
print_labeled_json("Result: local tool outputs", fallback_orders)
print_labeled_text("Result: final model answer", parallel_answer)
print_response_summary(parallel_final)
print_key_takeaway('Local lookup outputs keep the parallel-tool pattern understandable if not every call is returned.')
except Exception as exc:
handle_example_error("Parallel tool calls", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field
value
model
openai.gpt-5.4
max_output_tokens
1024
store
False
parallel_tool_calls
True
tools
get_order_status, get_customer_profile
tool_choice
auto
input
1 item(s): user: Use get_order_status for ORDER-8831 and ORDER-2044, then summarize whether Maya has one shipping problem or multiple active shipping problems in two labeled plain-text lines. Do not use leading hyphens or bold text.
[
{
"order_id": "ORDER-8831",
"customer_id": "CUST-1042",
"item": "standing desk replacement",
"status": "delayed",
"carrier_scan": "No movement for 36 hours at Denver sort center",
"promised_delivery": "2026-06-01",
"recommended_policy": "If delay exceeds 48 hours, offer expedited replacement or 15% concession with agent approval."
},
{
"order_id": "ORDER-2044",
"customer_id": "CUST-1042",
"item": "ergonomic chair",
"status": "delivered",
"carrier_scan": "Delivered yesterday at front desk",
"promised_delivery": "2026-05-29",
"recommended_policy": "Confirm delivery details before opening a replacement request."
}
]
<IPython.core.display.HTML object>
Resultado: respuesta final del modelo
Order statuses: ORDER-8831 is delayed and ORDER-2044 was delivered yesterday.
Shipping problems: Maya has one active shipping problem, because only ORDER-8831 is currently delayed while ORDER-2044 is already delivered.
Conclusión clave: Las salidas de búsqueda locales hacen que el patrón de herramientas paralelas sea comprensible, incluso si no se devuelve cada llamada.
4.4 Usar una herramienta de texto personalizada
Las herramientas personalizadas pasan texto de formato libre a la lógica propiedad de la aplicación en lugar de requerir un objeto de argumento JSON estructurado. Esta celda define un normalizador de notas de soporte, solicita una llamada a una herramienta personalizada e incluye un fallback local si el endpoint devuelve texto ordinario en lugar de una llamada personalizada. Inspecciona los tipos de elementos de salida y la nota normalizada.
from __future__ import annotations
custom_tools = [
{
"type": "custom",
"name": "normalize_support_note",
"description": "Normalize a freeform support note written by an agent. Input is plain text.",
"format": {"type": "text"},
}
]
def normalize_support_note_text(note: str) -> str:
fields = [part.strip().upper() for part in note.split("|")]
labels = ["ORDER_ID", "CUSTOMER_ID", "ISSUE", "CUSTOMER_REQUEST", "POLICY_OPTION"]
return "\n".join(
f"{label}: {value}"
for label, value in zip(labels, fields)
if value
)
support_note = "order-8831 | cust-1042 | replacement delayed | customer wants supervisor | offer expedited replacement or 15% concession"
custom_input = [{
"role": "user",
"content": (
"Call normalize_support_note with this exact note. Do not answer directly; "
f"send the note to the custom tool: {support_note}"
),
}]
custom_request = {
"model": MODEL_ID,
"input": custom_input,
"tools": custom_tools,
"tool_choice": {"type": "custom", "name": "normalize_support_note"},
"max_output_tokens": 1024,
"store": False,
}
print_request_shape(custom_request)
print_labeled_text("Result: local fallback normalization", normalize_support_note_text(support_note))
try:
custom_plan = create_response(**custom_request)
returned_item_types = [item.get("type") for item in response_items(custom_plan)]
try:
custom_call = first_output_item(custom_plan, "custom_tool_call")
if custom_call is None:
raise LookupError("No custom_tool_call item returned.")
tool_input = custom_call.get("input", "").strip()
normalized_note = normalize_support_note_text(tool_input)
record_check("Custom tools", "pass", {"output_item_types": returned_item_types, "normalized_note": normalized_note})
record_response("Normalized support note", "text", normalized_note)
print_labeled_json("Result: returned output item types", returned_item_types)
print_labeled_text("Result: custom tool input", tool_input)
print_labeled_text("Result: application-owned normalized output", normalized_note)
except LookupError:
fallback_text = output_text(custom_plan).strip() or "No text content was returned."
normalized_note = normalize_support_note_text(support_note)
record_check("Custom tools", "warn", {
"expected": "custom_tool_call item named normalize_support_note",
"actual_output_item_types": returned_item_types,
"meaning": "The model response did not include a custom-tool invocation, so the application fallback normalization is shown for teaching.",
})
record_response("Custom tool text fallback", "text", fallback_text)
record_response("Application-owned normalization fallback", "text", normalized_note)
print_labeled_json("Result: returned output item types", returned_item_types or ["no typed output items returned"])
print_labeled_text("Result: model text response", fallback_text)
print_labeled_text("Result: application-owned normalization", normalized_note)
print_response_summary(custom_plan)
print_key_takeaway('Custom tools are useful when the application owns a freeform parsing or execution step.')
except Exception as exc:
handle_example_error("Custom tools", exc)
1 item(s): user: Call normalize_support_note with this exact note. Do not answer directly; send the note to the custom tool: order-8831 | cust-1042 | replacement delayed | customer wants supervisor | offer expedited replacement or 15% co...
Conclusión clave: Las herramientas personalizadas son útiles cuando la aplicación posee un paso de análisis o ejecución de formato libre.
5. Enviar entrada de archivo directa
La entrada de archivo directa es independiente de las herramientas gestionadas por la aplicación. Se puede incluir un archivo en la solicitud de Responses actual como un elemento input_file junto con instrucciones de texto, lo cual es útil cuando el modelo debe leer el archivo para este turno sin configurar un índice de recuperación.
5.1 Adjuntar un PDF como input_file
Esta celda genera una pequeña transcripción PDF en memoria, la adjunta como datos de archivo base64 y solicita campos JSON exactos del documento. Inspecciona la vista previa del PDF, los campos esperados, la respuesta analizada y el resumen de uso.
from __future__ import annotations
def make_simple_pdf(lines: list[str]) -> bytes:
def pdf_escape(text: str) -> str:
return text.replace("\\", "\\\\").replace("(", "\\(").replace(")", "\\)")
stream_lines = ["BT", "/F1 11 Tf", "72 740 Td", "15 TL"]
for idx, line in enumerate(lines):
if idx:
stream_lines.append("T*")
stream_lines.append(f"({pdf_escape(line)}) Tj")
stream_lines.append("ET")
stream = "\n".join(stream_lines).encode("latin-1", "replace")
objects = [
b"<< /Type /Catalog /Pages 2 0 R >>",
b"<< /Type /Pages /Kids [3 0 R] /Count 1 >>",
b"<< /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] /Resources << /Font << /F1 4 0 R >> >> /Contents 5 0 R >>",
b"<< /Type /Font /Subtype /Type1 /BaseFont /Helvetica >>",
b"<< /Length " + builtins.str(len(stream)).encode("ascii") + b" >>\nstream\n" + stream + b"\nendstream",
]
pdf = b"%PDF-1.4\n"
offsets = [0]
for idx, obj in enumerate(objects, start=1):
offsets.append(len(pdf))
pdf += f"{idx} 0 obj\n".encode("ascii") + obj + b"\nendobj\n"
xref_offset = len(pdf)
pdf += f"xref\n0 {len(objects) + 1}\n0000000000 65535 f \n".encode("ascii")
for offset in offsets[1:]:
pdf += f"{offset:010d} 00000 n \n".encode("ascii")
pdf += f"trailer\n<< /Size {len(objects) + 1} /Root 1 0 R >>\nstartxref\n{xref_offset}\n%%EOF\n".encode("ascii")
return pdf
file_lines = [
"BrightCart support transcript",
"Ticket: TICKET-7429",
"Customer: Maya Chen",
"Order: ORDER-8831",
"Product: Standing desk replacement",
"Issue: Replacement for a damaged item is delayed and carrier scan has not moved",
"Customer request: Supervisor callback and refund options",
"Policy options: expedited replacement or 15% concession with agent approval after 48-hour delay",
]
file_text = "\n".join(file_lines)
pdf_data = base64.b64encode(make_simple_pdf(file_lines)).decode("utf-8")
expected_direct_file_fields = {
"ticket_id": "TICKET-7429",
"customer": "Maya Chen",
"order_id": "ORDER-8831",
"product": "Standing desk replacement",
}
direct_file_request = {
"model": MODEL_ID,
"input": [
{
"role": "user",
"content": [
{
"type": "input_file",
"filename": "brightcart-support-transcript.pdf",
"file_data": f"data:application/pdf;base64,{pdf_data}",
},
{
"type": "input_text",
"text": (
"Read the attached PDF support transcript and return JSON with keys "
"ticket_id, customer, order_id, product, issue, requested_resolution, and policy_options. "
"Use exact values from the file. Do not return null for fields that are present in the file."
),
},
],
}
],
"text": {"format": {"type": "json_object"}},
"max_output_tokens": 1024,
"store": False,
}
print_labeled_text("Result: PDF transcript preview", file_text)
print_request_shape(direct_file_request)
print_labeled_json("Result: expected fields", expected_direct_file_fields)
try:
direct_file_response = create_response(**direct_file_request)
raw_direct_file_output = output_text(direct_file_response).strip()
try:
direct_file_payload = json.loads(raw_direct_file_output)
missing_or_empty = [
key for key, expected in expected_direct_file_fields.items()
if builtins.str(direct_file_payload.get(key, "")).strip().lower() != expected.lower()
]
null_fields = [key for key, value in direct_file_payload.items() if value in {None, "", []}]
if missing_or_empty or null_fields:
record_check("Direct file inputs", "warn", {
"message": "The request completed, but the model did not extract the expected values from the attached PDF.",
"missing_or_unexpected_fields": missing_or_empty,
"empty_fields": null_fields,
"payload": direct_file_payload,
})
record_response("Support transcript extraction returned by model", "json", direct_file_payload)
print_labeled_text("Result", "The request completed, but the model did not extract the expected values from the attached PDF.")
print_labeled_json("Result: returned JSON", direct_file_payload)
else:
record_check("Direct file inputs", "pass", direct_file_payload)
record_response("Support transcript extraction", "json", direct_file_payload)
print_labeled_json("Result", direct_file_payload)
except Exception as parse_exc:
record_check("Direct file inputs", "warn", {
"message": "The request completed, but the response was not valid JSON.",
"text_sample": raw_direct_file_output[:600],
"error": builtins.str(parse_exc),
})
record_response("Support transcript extraction text", "text", raw_direct_file_output[:1200])
print_labeled_text("Result", raw_direct_file_output[:1200])
print_response_summary(direct_file_response)
print_key_takeaway('Direct file input is useful when the file should be read in the current request context.')
except Exception as exc:
handle_example_error("Direct file inputs", exc)
<IPython.core.display.HTML object>
Resultado: vista previa de la transcripción PDF
BrightCart support transcript
Ticket: TICKET-7429
Customer: Maya Chen
Order: ORDER-8831
Product: Standing desk replacement
Issue: Replacement for a damaged item is delayed and carrier scan has not moved
Customer request: Supervisor callback and refund options
Policy options: expedited replacement or 15% concession with agent approval after 48-hour delay
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field
value
model
openai.gpt-5.4
max_output_tokens
1024
store
False
text format
json_object
input
1 item(s): user: input_file: brightcart-support-transcript.pdf; input_text: Read the attached PDF support transcript and return JSON with keys ticket_id, customer, order_id, product, issue, reques...
{"ticket_id":"TICKET-7429","customer":"Maya Chen","order_id":"ORDER-8831","product":"Standing desk replacement","issue":"Replacement for a damaged item is delayed and carrier scan has not moved","requested_resolution":"Supervisor callback and refund options","policy_options":"expedited replacement or 15% concession with agent approval after 48-hour delay"}
Conclusión clave: La entrada de archivo directa es útil cuando el archivo debe leerse en el contexto de la solicitud actual.
6. Gestionar el estado de la conversación
El estado de la conversación determina cómo los turnos de seguimiento reciben el contexto anterior. La API de Responses admite la continuación almacenada con previous_response_id, y las aplicaciones también pueden gestionar el estado por sí mismas reenviando el historial de entrada relevante. Esta sección compara ambos patrones y luego muestra el contexto de razonamiento cifrado donde se admite.
6.1 Continuar con previous_response_id
Usa previous_response_id para continuar desde una respuesta almacenada sin reenviar el prompt completo anterior. La primera solicitud almacena los detalles del caso BrightCart; la segunda solicitud pasa solo la nueva instrucción de seguimiento más el ID de respuesta anterior. Inspecciona si el seguimiento conserva el pedido, el cliente, el problema y la siguiente acción.
from __future__ import annotations
promised_delivery = (date.today() + timedelta(days=2)).isoformat()
stateful_seed_input = (
f"Customer Maya Chen opened ticket TICKET-4812 about order ORDER-8831. "
"The item is a standing desk replacement for a damaged delivery. "
f"The promised delivery date is {promised_delivery}, but the carrier scan has not moved in 36 hours. "
"Customer sentiment is frustrated because this is the second attempt. "
"Support policy says to offer expedited replacement or a 15% concession if the delay exceeds 48 hours. "
"Escalation owner is Tier 2 Returns."
)
stateful_followup_input = "Return five labeled lines: ticket ID, order ID, customer name, issue, and next best action."
stateful_request_shape = {
"model": MODEL_ID,
"input": stateful_followup_input,
"previous_response_id": "<response-id-from-prior-stored-turn>",
"max_output_tokens": 1024,
"store": False,
}
print_request_shape(stateful_request_shape)
try:
stateful_turn_1 = create_response(model=MODEL_ID, input=stateful_seed_input, max_output_tokens=1024, store=True)
remember_stored_response(stateful_turn_1)
stateful_turn_2 = create_response(model=MODEL_ID, input=stateful_followup_input, previous_response_id=stateful_turn_1.id, max_output_tokens=1024, store=False)
text = output_text(stateful_turn_2).strip()
require("order-8831" in text.lower() or "maya" in text.lower(), "Stateful continuation response missed expected support context.")
record_check("Stateful continuation", "pass", stateful_turn_1.id)
record_response("Stateful support handoff", "text", text)
print_labeled_text("Result", text)
print_response_summary(stateful_turn_2)
print_key_takeaway('previous_response_id lets a follow-up use stored context without resending the full prior turn.')
except Exception as exc:
handle_example_error("Stateful continuation", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field
value
model
openai.gpt-5.4
max_output_tokens
1024
store
False
previous_response_id
<response-id-from-prior-stored-turn>
input
Return five labeled lines: ticket ID, order ID, customer name, issue, and next best action.
<IPython.core.display.HTML object>
Resultado
Ticket ID: TICKET-4812
Order ID: ORDER-8831
Customer Name: Maya Chen
Issue: Replacement standing desk shipment for damaged delivery has had no carrier movement for 36 hours; customer is frustrated because this is the second attempt
Next Best Action: Monitor until 48 hours without movement, then offer expedited replacement or 15% concession and escalate to Tier 2 Returns if needed
Conclusión clave: previous_response_id permite que un seguimiento use el contexto almacenado sin reenviar el turno anterior completo.
6.2 Reconstruir el contexto sin estado
La continuación sin estado significa que la aplicación envía el historial relevante en cada solicitud. Esto es adecuado cuando tu producto ya posee el almacenamiento de conversaciones, la política de retención o los requisitos de auditoría. Esta celda envía un historial de chat corto más una nueva instrucción de traspaso e inspecciona el resumen y el uso de tokens.
from __future__ import annotations
stateless_history = [
{"role": "user", "content": "Support chat TICKET-3920: Customer Jordan Lee says ORDER-7718 arrived with a cracked monitor stand."},
{"role": "assistant", "content": "Captured damaged-item issue for ORDER-7718 and asked for preferred resolution."},
{"role": "user", "content": "Jordan wants a replacement shipped this week and asks whether the damaged item must be returned first."},
]
stateless_payload = {
"model": MODEL_ID,
"input": stateless_history + [{"role": "user", "content": "Summarize this support chat for the next agent in five labeled plain-text lines. Do not use leading hyphens or bold text."}],
"max_output_tokens": 1024,
"store": False,
}
print_request_shape(stateless_payload)
try:
stateless_response = create_response(**stateless_payload)
stateless_text = output_text(stateless_response).strip()
require(stateless_text, "Stateless continuation response did not return text.")
record_check("Stateless continuation", "pass", summarize_response(stateless_response))
record_response("Stateless support handoff", "text", stateless_text)
print_labeled_text("Result", stateless_text)
print_response_summary(stateless_response)
print_key_takeaway('Stateless continuation sends the relevant history with each request when the application owns conversation storage.')
except Exception as exc:
handle_example_error("Stateless continuation", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field
value
model
openai.gpt-5.4
max_output_tokens
1024
store
False
input
4 item(s): user: Support chat TICKET-3920: Customer Jordan Lee says ORDER-7718 arrived with a cracked monitor stand.; assistant: Captured damaged-item issue for ORDER-7718 and asked for preferred resolution.; user: Jordan wants a replacement shipped this week and asks whether the damaged item must be returned first.; user: Summarize this support chat for the next agent in five labeled plain-text lines. Do not use leading hyphens or bold text.
<IPython.core.display.HTML object>
Resultado
Customer: Jordan Lee reported ORDER-7718 arrived with a cracked monitor stand.
Issue: Damaged item; monitor stand is cracked on arrival.
Requested Resolution: Customer wants a replacement shipped this week.
Open Question: Jordan asked whether the damaged item must be returned before replacement is sent.
Status: Damage claim captured and awaiting next-agent confirmation on replacement timing and return requirement.
Conclusión clave: La continuación sin estado envía el historial relevante con cada solicitud cuando la aplicación posee el almacenamiento de conversaciones.
6.3 Llevar el contexto de razonamiento cifrado
Los modelos con capacidad de razonamiento pueden devolver elementos de razonamiento y contenido de razonamiento cifrado cuando se les solicita. Esta celda solicita metadatos de razonamiento cifrados, lleva elementos de respuesta anteriores a una solicitud de seguimiento e inspecciona si se devolvió contenido cifrado. El texto de razonamiento oculto no se expone; la aplicación solo lleva el contexto opaco hacia adelante cuando es compatible.
Documentos oficiales: Modelos de razonamiento describe los modelos de razonamiento y el esfuerzo de razonamiento en los flujos de trabajo de respuestas.
from __future__ import annotations
encrypted_history = [
{"role": "user", "content": "For a customer-support assistant handling names, order IDs, and refund context, compare stateful and stateless continuation in two sentences."}
]
encrypted_turn_payload = {
"model": MODEL_ID,
"input": encrypted_history,
"reasoning": {"effort": "medium"},
"include": ["reasoning.encrypted_content"],
"max_output_tokens": 1024,
"store": False,
}
print_request_shape(encrypted_turn_payload)
try:
encrypted_turn_1 = create_response(**encrypted_turn_payload)
encrypted_turn_2 = create_response(
model=MODEL_ID,
input=encrypted_history + response_items(encrypted_turn_1) + [
{"role": "user", "content": "Based on the prior reasoning context, recommend one approach for a regulated support workflow in two labeled plain-text lines. Do not use leading hyphens or bold text."}
],
max_output_tokens=1024,
store=False,
)
reasoning_items = [item for item in response_items(encrypted_turn_1) if item.get("type") == "reasoning"]
has_encrypted_content = any(item.get("encrypted_content") for item in reasoning_items)
record_check("Encrypted reasoning", "pass", {"encrypted_content_returned": has_encrypted_content, "reasoning_item_count": len(reasoning_items)})
encrypted_answer = output_text(encrypted_turn_2).strip()
record_response("State strategy recommendation", "text", encrypted_answer)
print_labeled_json("Result: reasoning metadata", {
"returned_item_types": [item.get("type") for item in response_items(encrypted_turn_1)],
"encrypted_reasoning_content_returned": has_encrypted_content,
})
print_labeled_text("Result: follow-up answer", encrypted_answer)
print_response_summary(encrypted_turn_2)
print_key_takeaway('Encrypted reasoning content can be carried forward where supported without exposing hidden reasoning text.')
except Exception as exc:
handle_example_error("Encrypted reasoning", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field
value
model
openai.gpt-5.4
max_output_tokens
1024
store
False
reasoning
{'effort': 'medium'}
include
['reasoning.encrypted_content']
input
1 item(s): user: Para un asistente de atención al cliente que maneja nombres, ID de pedidos y contexto de reembolsos, compara la continuación con estado y sin estado en dos oraciones.
Recomendación: Continuación sin estado
Razón: En un flujo de trabajo de soporte regulado, requerir que los nombres, ID de pedidos y el contexto de reembolso se proporcionen explícitamente en cada turno mejora la controlabilidad, la auditabilidad y la minimización de datos, reduciendo el riesgo de retención no intencionada o fuga entre sesiones.
Conclusión clave: El contenido de razonamiento cifrado se puede llevar adelante donde sea compatible sin exponer el texto de razonamiento oculto.
7. Usar el almacenamiento en caché de prompts
El almacenamiento en caché de prompts mejora la latencia y el costo cuando las solicitudes comparten un prefijo estático exacto.
7.1 Comparar dos solicitudes con clave de caché
Esta celda coloca texto de política estable de BrightCart al principio de la entrada, envía la misma solicitud dos veces con un prompt_cache_key y compara los metadatos de los tokens. Inspecciona cached_input_tokens en la segunda respuesta cuando el endpoint devuelve los detalles de la caché.
Nota: PROMPT_CACHE_RETENTION se selecciona del MODEL_ID activo. Usa 24h para openai.gpt-5.5 y modelos posteriores, y in_memory para openai.gpt-5.4.
from __future__ import annotations
base_support_policy = [
"BrightCart support policy:",
"1. Be empathetic, concise, and specific about the customer's order.",
"2. Do not promise refunds, credits, or delivery dates unless the policy context supports it.",
"3. For damaged-item replacements, check replacement status before offering concessions.",
"4. If a replacement delay exceeds 48 hours, offer expedited replacement or a 15% concession subject to agent approval.",
]
policy_reference_paragraph = (
"Expanded cacheable policy context: BrightCart agents should identify the customer, order ID, replacement status, "
"carrier scan age, promised delivery window, item category, prior concessions, and supervisor approval needs before "
"drafting a customer-facing answer. The assistant should preserve a calm tone, avoid unsupported promises, separate "
"confirmed facts from assumptions, recommend one clear next action, and document why any escalation, expedited "
"replacement, or concession is appropriate. Repeated policy context like this is intentionally stable across many "
"requests so prompt caching can reuse the prefix when the same cache key is supplied."
)
expanded_policy_context = "\n".join(
f"Policy reference paragraph {idx + 1}: {policy_reference_paragraph}"
for idx in range(32)
)
stable_support_policy = "\n".join(base_support_policy + [expanded_policy_context])
cache_input = [
{"role": "system", "content": stable_support_policy},
{"role": "user", "content": "Draft a two-sentence agent reply for Maya Chen about delayed replacement order ORDER-8831."},
]
estimated_cache_input_words = len(json.dumps(cache_input).split())
require(estimated_cache_input_words > 2048, f"Prompt-cache input should be over 2048 words; found {estimated_cache_input_words}.")
cache_payload = {
"model": MODEL_ID,
"input": cache_input,
"prompt_cache_key": "brightcart-support-policy-v1",
"prompt_cache_retention": PROMPT_CACHE_RETENTION,
"max_output_tokens": 1024,
"store": False,
}
print_request_shape(cache_payload)
print_labeled_json("Prompt-cache input size", {"estimated_input_words": estimated_cache_input_words, "target_minimum_tokens": 2048})
try:
cache_response_1 = create_response(**cache_payload)
cache_response_2 = create_response(**cache_payload)
cache_summary_1 = summarize_response(cache_response_1)
cache_summary_2 = summarize_response(cache_response_2)
cache_comparison = pd.DataFrame([
{
"request": "first",
"input_tokens": cache_summary_1.get("input_tokens"),
"cached_input_tokens": cache_summary_1.get("cached_input_tokens"),
"output_tokens": cache_summary_1.get("output_tokens"),
"total_tokens": cache_summary_1.get("total_tokens"),
},
{
"request": "second",
"input_tokens": cache_summary_2.get("input_tokens"),
"cached_input_tokens": cache_summary_2.get("cached_input_tokens"),
"output_tokens": cache_summary_2.get("output_tokens"),
"total_tokens": cache_summary_2.get("total_tokens"),
},
])
record_check("Prompt caching", "pass" if cache_summary_2.get("cached_input_tokens") is not None else "warn", {"first": cache_summary_1, "second": cache_summary_2})
cache_reply = output_text(cache_response_2).strip()
record_response("Prompt-cache token comparison", "table", cache_comparison)
record_response("Cached support-policy reply", "text", cache_reply)
print_labeled_text("Result", cache_reply)
print_labeled_json("First request summary", cache_summary_1)
print_labeled_json("Second request summary", cache_summary_2)
print_label("Response summary")
display_wrapped_table(cache_comparison, max_col_width_px=260)
print_key_takeaway('cached_input_tokens is the metadata field to inspect for prompt-cache reuse.')
except Exception as exc:
handle_example_error("Prompt caching", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field
value
model
openai.gpt-5.4
max_output_tokens
1024
store
False
prompt_cache_key
brightcart-support-policy-v1
prompt_cache_retention
in_memory
input
2 item(s): system: Política de soporte de BrightCart: 1. Sé empático, conciso y específico sobre el pedido del cliente. 2. No prometas reembolsos, créditos o fechas de entrega a menos que el contexto de la política lo respalde. 3. Para reemplazos de artículos dañados...; user: Redacta una respuesta de agente de dos oraciones para Maya Chen sobre el pedido de reemplazo retrasado ORDER-8831.
Hola Maya, lamento que tu pedido de reemplazo ORDER-8831 esté retrasado. Estoy verificando el estado más reciente del reemplazo y del transportista para poder confirmar el mejor siguiente paso para ti sin hacerte esperar más de lo necesario.
Conclusión clave: cached_input_tokens es el campo de metadatos a inspeccionar para la reutilización de la caché de prompts.
8. Ejecutar trabajo en segundo plano
El modo en segundo plano inicia una respuesta de forma asíncrona y permite que la aplicación consulte el estado final.
8.1 Enviar y consultar una respuesta en segundo plano
Esta celda envía background=true, almacena el ID de la respuesta, consulta mientras el estado está en cola o en progreso, y luego imprime el resumen final del gerente. Inspecciona el historial de estado, el estado final, el ID de la respuesta y el resumen de tokens.
from __future__ import annotations
backlog = """
Same-day BrightCart support backlog:
1. 18 delayed-order contacts, mostly from the West Coast distribution lane.
2. 7 damaged-item replacement contacts; 3 mention replacement delays.
3. 5 return-window exception requests after holiday promotions.
""".strip()
background_payload = {
"model": MODEL_ID,
"input": f"Return exactly three labeled plain-text lines for a support-manager summary: theme, risk, next action. Keep each line under 12 words. Do not use leading hyphens or bold text.\n\n{backlog}",
"background": True,
"max_output_tokens": 1024,
"store": True,
}
print_request_shape(background_payload)
try:
background_response = create_response(**background_payload)
remember_stored_response(background_response)
status_history = [getattr(background_response, "status", None)]
for _ in range(15):
if getattr(background_response, "status", None) not in {"queued", "in_progress"}:
break
time.sleep(2)
background_response = retrieve_response(background_response.id)
status_history.append(getattr(background_response, "status", None))
background_summary = summarize_response(background_response)
manager_summary = output_text(background_response).strip()
require(manager_summary, "Background response did not return text.")
status = "pass" if background_summary.get("status") in {None, "completed"} else "warn"
record_check("Background mode", status, {"status_history": status_history, "id": getattr(background_response, "id", None), "final_status": background_summary.get("status")})
record_response("Background manager summary", "text", manager_summary)
print_labeled_json("Result: status history", status_history)
print_labeled_text("Result: manager summary", manager_summary)
print_response_summary(background_summary)
print_key_takeaway('Background mode starts work asynchronously and lets the application poll by response ID.')
except Exception as exc:
handle_example_error("Background mode", exc)
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field
value
model
openai.gpt-5.4
max_output_tokens
1024
store
True
background
True
input
Devuelve exactamente tres líneas de texto plano etiquetadas para un resumen del gerente de soporte: tema, riesgo, siguiente acción. Mantén cada línea por debajo de 12 palabras. No uses guiones iniciales ni texto en negrita. Retraso de soporte de BrightCart del mismo día: 1. 18 contactos de pedidos retrasados, la mayoría de la We...
<IPython.core.display.HTML object>
Resultado: historial de estado
[
"in_progress",
"completed"
]
<IPython.core.display.HTML object>
Resultado: resumen del gerente
tema: Los retrasos en el envío dominan, especialmente en la ruta de distribución de la Costa Oeste.
riesgo: Aumento de la insatisfacción por retrasos, reemplazos y excepciones de devolución.
siguiente acción: Escalar los problemas de la ruta de la Costa Oeste y revisar la política de devoluciones de vacaciones.
Conclusión clave: El modo en segundo plano inicia el trabajo de forma asíncrona y permite que la aplicación consulte por ID de respuesta.
9. Compactar el contexto de larga duración
La compactación reduce el estado de una conversación larga a hechos duraderos, preguntas abiertas, restricciones y próximas acciones. Esta celda documenta el patrón de compactación del lado de la aplicación como un pequeño objeto JSON para que el concepto sea claro sin agregar otra ruta de función en vivo. Inspecciona qué hechos se mantienen y qué detalles se omiten antes del siguiente turno.
from __future__ import annotations
compaction_note = {
"feature": "Compaction",
"how_to_apply": "Summarize older support turns into durable facts, open questions, policy constraints, and next actions before continuing the workflow.",
"brightcart_example": {
"durable_facts": ["Customer Maya Chen", "ORDER-8831", "replacement delayed", "carrier scan stale"],
"policy_constraints": ["Do not promise refund without eligibility", "Offer expedited replacement or 15% concession after 48-hour delay with approval"],
"next_action": "Check latest carrier scan and supervisor callback status.",
},
}
record_check("Compaction", "documented", compaction_note)
record_response("Compacted support context", "json", compaction_note)
print_json(compaction_note)
<IPython.core.display.HTML object>
JSON
{
"feature": "Compaction",
"how_to_apply": "Resume los turnos de soporte anteriores en hechos duraderos, preguntas abiertas, restricciones de política y próximas acciones antes de continuar el flujo de trabajo.",
"brightcart_example": {
"durable_facts": [
"Cliente Maya Chen",
"ORDER-8831",
"reemplazo retrasado",
"escaneo del transportista obsoleto"
],
"policy_constraints": [
"No prometas reembolso sin elegibilidad",
"Ofrece reemplazo expedito o concesión del 15% después de un retraso de 48 horas con aprobación"
],
"next_action": "Verifica el último escaneo del transportista y el estado de la llamada del supervisor."
}
}
Las verificaciones operativas rápidas son comprobaciones de configuración ligeras, no una prueba de carga o una medición a nivel de servicio. Esta celda envía tres solicitudes cortas, mide el tiempo transcurrido local, resume la tasa de éxito y el uso de tokens, e infiere la región de la URL base de Bedrock configurada. Inspecciona la latencia, el estado de finalización, las salidas de muestra y los totales de tokens.
from __future__ import annotations
def infer_region_from_base_url(base_url: str) -> str | None:
host = normalize_base_url(base_url).replace("https://", "").split("/")[0]
for part in host.split("."):
if part.count("-") >= 2 and any(char.isdigit() for char in part):
return part
return None
def percentile(values: list[float], pct: float) -> float | None:
if not values:
return None
ordered = sorted(values)
index = min(len(ordered) - 1, max(0, round((pct / 100) * (len(ordered) - 1))))
return round(ordered[index], 3)
operations_features = ["Latency runtime example", "Throughput runtime example", "Reliability runtime example", "Region check"]
operations_payload = {"model": MODEL_ID, "input": "Reply with one short customer-support sentence.", "service_tier": "auto", "max_output_tokens": 1024, "store": False}
print_request_shape(operations_payload)
if not RUN_RESPONSIVENESS_CHECK:
record_check("Endpoint responsiveness", "skipped", "BEDROCK_RESPONSIVENESS_CHECK is disabled.")
print_labeled_text("Result", "Responsiveness check disabled.")
else:
prompts = [
"Reply in one short sentence: apologize for a delayed replacement order.",
"Reply with one metric name for support-assistant quality.",
"Reply in one short sentence: hand off a return exception to a supervisor.",
]
samples = []
for idx, prompt in enumerate(prompts):
started = time.perf_counter()
try:
response = create_response(model=MODEL_ID, input=prompt, service_tier="auto", max_output_tokens=1024, store=False)
elapsed = time.perf_counter() - started
summary = summarize_response(response)
text = output_text(response).strip()
samples.append({
"ok": bool(text),
"latency_seconds": round(elapsed, 3),
"output_tokens": summary.get("output_tokens") or 0,
"total_tokens": summary.get("total_tokens") or 0,
"status": summary.get("status"),
"sample_output": text[:140],
})
except Exception as exc:
elapsed = time.perf_counter() - started
samples.append({"ok": False, "latency_seconds": round(elapsed, 3), "error": describe_api_error(exc)})
successes = [sample for sample in samples if sample["ok"]]
completed = [sample for sample in successes if sample.get("status") in {None, "completed"}]
latencies = [sample["latency_seconds"] for sample in successes]
responsiveness_summary = {
"region_hint": infer_region_from_base_url(BASE_URL),
"base_url_host": normalize_base_url(BASE_URL).replace("https://", "").split("/")[0],
"sample_count": len(samples),
"success_rate": len(successes) / len(samples) if samples else 0,
"completed_rate": len(completed) / len(samples) if samples else 0,
"avg_latency_seconds": round(sum(latencies) / len(latencies), 3) if latencies else None,
"p50_latency_seconds": percentile(latencies, 50),
"p90_latency_seconds": percentile(latencies, 90),
"total_output_tokens": sum(sample.get("output_tokens", 0) for sample in samples),
"total_tokens": sum(sample.get("total_tokens", 0) for sample in samples),
}
status = "pass" if len(successes) == len(samples) and len(completed) == len(samples) else "warn"
for feature in operations_features:
record_check(feature, status, responsiveness_summary)
record_response("Endpoint responsiveness summary", "json", {**responsiveness_summary, "samples": samples})
print_labeled_json("Result", responsiveness_summary)
print_label("Response summary")
display_wrapped_table(pd.DataFrame(samples), max_col_width_px=360)
print_key_takeaway('Responsiveness samples are setup checks, not a load test or service-level measurement.')
<IPython.core.display.HTML object>
Forma de la solicitud
<IPython.core.display.HTML object>
field
value
model
openai.gpt-5.4
max_output_tokens
1024
store
False
service_tier
auto
input
Responde con una oración corta de atención al cliente.
Estoy escalando la excepción de devolución a un supervisor.
<IPython.core.display.HTML object>
Conclusión clave: Las muestras de capacidad de respuesta son verificaciones de configuración, no una prueba de carga o una medición a nivel de servicio.
11. Limpia y revisa los resultados
Las respuestas almacenadas creadas por los ejemplos de ciclo de vida, continuación con estado y segundo plano se rastrean en STORED_RESPONSE_IDS. Esta celda final intenta eliminar las respuestas almacenadas cuando la limpieza está habilitada, luego imprime el resumen de la ejecución y la galería de respuestas de ejemplo. Primero, inspecciona las advertencias; estas suelen identificar diferencias en la configuración del endpoint, la disponibilidad del modelo o el soporte de características.
from __future__ import annotations
cleanup_rows = []
tracked_ids = list(dict.fromkeys(STORED_RESPONSE_IDS))
if not tracked_ids:
cleanup_rows.append({"response_id": "none", "status": "no stored responses tracked", "detail": ""})
elif not CLEAN_UP_STORED_RESPONSES:
for stored_id in tracked_ids:
cleanup_rows.append({"response_id": stored_id, "status": "skipped", "detail": "BEDROCK_CLEANUP_STORED_RESPONSES is disabled"})
record_check("Stored response cleanup", "skipped", cleanup_rows)
else:
for stored_id in tracked_ids:
try:
delete_result = delete_response(stored_id)
cleanup_rows.append({"response_id": stored_id, "status": "deleted", "detail": compact_text(to_dict(delete_result), 240)})
except Exception as exc:
cleanup_rows.append({"response_id": stored_id, "status": "warn", "detail": describe_api_error(exc)})
cleanup_status = "pass" if all(row["status"] == "deleted" for row in cleanup_rows) else "warn"
record_check("Stored response cleanup", cleanup_status, cleanup_rows)
print_label("Stored response cleanup")
display_wrapped_table(pd.DataFrame(cleanup_rows), max_col_width_px=520)
summary_df = pd.DataFrame(RESULTS_SUMMARY)
print_label("Run summary")
display_wrapped_table(summary_df, max_col_width_px=620)
print_label("Example responses")
print_response_gallery()
{"ticket_id": "TICKET-7429", "category": "delivery_delay", "priority": "urgent", "customer_sentiment": "frustrated and time-sensitive", "summary": "Customer Maya Chen reports that ORDER-8831 is a replacement shipment for a previously damaged standing desk. The replacement is now 2 days late, carrier tracking has not updated, and she needs the desk delivered before Monday. She is requesting a supervisor callback and wants to know refund options if the replacement cannot arrive in time.", "require...
JSON mode
pass
{"customer_name": "Maya Chen", "order_id": "ORDER-8831", "issue_summary": "Customer is asking about a delayed replacement order. The carrier tracking scan is stale and has not updated.", "next_step": "Handoff to support to investigate the carrier delay, verify shipment status, and provide Maya Chen with an update or resolution.", "metrics_to_watch": ["tracking_scan_recency", "carrier_exception_status", "replacement_order_delivery_eta", "customer_follow_up_time"]}
{"message": "The request completed, but the response was not valid JSON.", "text_sample": "{\"ticket_id\":\"TICKET-7429\",\"customer\":\"Maya Chen\",\"order_id\":\"ORDER-8831\",\"product\":\"Standing desk replacement\",\"issue\":\"Replacement for a damaged item is delayed and carrier scan has not moved\",\"requested_resolution\":\"Supervisor callback and refund options\",\"policy_options\":\"expedited replacement or 15% concession with agent approval after 48-hour delay\"}", "error": "unhashable...
{"feature": "Compaction", "how_to_apply": "Summarize older support turns into durable facts, open questions, policy constraints, and next actions before continuing the workflow.", "brightcart_example": {"durable_facts": ["Customer Maya Chen", "ORDER-8831", "replacement delayed", "carrier scan stale"], "policy_constraints": ["Do not promise refund without eligibility", "Offer expedited replacement or 15% concession after 48-hour delay with approval"], "next_action": "Check latest carrier scan and...
[{"response_id": "resp_cvhvh7y5ghwrpa35snvk4bzgcgthgxp4tgwkllmf5mrhs7dikfia", "status": "warn", "detail": {"exception_class": "AuthenticationError", "status_code": 401, "retryable": false, "request_id": "req_gkni5zyr7lkjkz2vfiwvkev2qgxs76crcwz5whhjrdkma7up3yta", "message": "Error code: 401 - {'error': {'code': 'invalid_api_key', 'message': 'The security token included in the request is invalid.', 'param': None, 'type': 'permission_denied_error'}}"}}, {"response_id": "resp_vjrtvnakcgxjhnq5b7cj7rt...
<IPython.core.display.HTML object>
Respuestas de ejemplo
<IPython.core.display.HTML object>
example
response_type
response
Endpoint verification
text
ok
First raw HTTPS request
text
Empathy: I’m sorry, Maya — your replacement order ORDER-8831 is delayed because the carrier reported a temporary transit hold at the regional sorting facility.\nAction: We’re monitoring the shipment closely and will send you an updated delivery estimate within 24 hours; if there’s no movement by then, we’ll review the next replacement or refund options with you.
SDK text generation
text
Use the Responses API to build a BrightCart support assistant that can answer customer questions, summarize policies, and guide users through common workflows like order tracking, refunds, and account updates. Ground the assistant in BrightCart documentation and connect it to relevant backend tools or APIs so it can retrieve live order data, check account status, and provide accurate, context-aware support responses. Design the experience around clear system instructions, structured tool calling, and conversation state management so the assistant stays on-brand, reliable, and safe when handling customer issues.
Create and retrieve response
text
goal: Help support agents explain delayed replacement orders, set expectations, and suggest next steps.\ndata needed: Order ID, replacement order status, shipment/tracking events, delay reason, estimated ship/delivery date, customer contact history, inventory/backorder status, and applicable refund or reship policy.\nhuman-review rule: Escalate to a human if the delay exceeds policy thresholds, tracking is inconsistent or missing, the order appears lost, the customer is high-risk or highly upset, or any refund/reship exception is requested.
Service tier and prompt cache request
text
Latency benefit: Prompt caching lets the BrightCart support assistant reuse previously processed context, reducing response time for repeated or similar requests.\nConsistency benefit: Prompt caching helps the BrightCart support assistant return more uniform answers by reusing the same established prompt context across interactions.
Structured ticket triage
json
{\n "ticket_id": "TICKET-7429",\n "category": "delivery_delay",\n "priority": "urgent",\n "customer_sentiment": "frustrated and time-sensitive",\n "summary": "Customer Maya Chen reports that ORDER-8831 is a replacement shipment for a previously damaged standing desk. The replacement is now 2 days late, carrier tracking has not updated, and she needs the desk delivered before Monday. She is requesting a supervisor callback and wants to know refund options if the replacement cannot arrive in time.",\n "required_actions": [\n "Review ORDER-8831 shipment status and confirm last carrier scan/update.",\n "Contact carrier or open a trace/escalation for stalled tracking.",\n "Check expedited reshipment or alternative fulfillment options to meet the before-Monday deadline.",\n "Arrange supervisor callback per customer request.",\n "Review and communicate refund options, including refund\n...
JSON support handoff
json
{\n "customer_name": "Maya Chen",\n "order_id": "ORDER-8831",\n "issue_summary": "Customer is asking about a delayed replacement order. The carrier tracking scan is stale and has not updated.",\n "next_step": "Handoff to support to investigate the carrier delay, verify shipment status, and provide Maya Chen with an update or resolution.",\n "metrics_to_watch": [\n "tracking_scan_recency",\n "carrier_exception_status",\n "replacement_order_delivery_eta",\n "customer_follow_up_time"\n ]\n}
Compact policy guidance
text
BrightCart’s delayed-replacement policy lets customers keep using the original item until the replacement arrives, then return the defective product within the allowed return window.
Detailed policy guidance
text
1. BrightCart sends replacements after customers return the original item and warehouse receipt is confirmed.\n2. This delay prevents duplicate shipments, verifies eligibility, and reduces fraud or inventory errors.\n3. Agents should explain timelines clearly, offer return instructions, and reassure customers once receipt is logged.
Order-status tool answer
text
Status: ORDER-8831 is delayed; carrier shows no movement for 36 hours at the Denver sort center, with promised delivery on 2026-06-01.\nNext best action: Monitor until the 48-hour threshold; if no movement then, contact the customer and offer either an expedited replacement or a 15% concession with agent approval.
Parallel order lookup fallback answer
text
Order statuses: ORDER-8831 is delayed and ORDER-2044 was delivered yesterday.\nShipping problems: Maya has one active shipping problem, because only ORDER-8831 is currently delayed while ORDER-2044 is already delivered.
{"ticket_id":"TICKET-7429","customer":"Maya Chen","order_id":"ORDER-8831","product":"Standing desk replacement","issue":"Replacement for a damaged item is delayed and carrier scan has not moved","requested_resolution":"Supervisor callback and refund options","policy_options":"expedited replacement or 15% concession with agent approval after 48-hour delay"}
Stateful support handoff
text
Ticket ID: TICKET-4812\nOrder ID: ORDER-8831\nCustomer Name: Maya Chen\nIssue: Replacement standing desk shipment for damaged delivery has had no carrier movement for 36 hours; customer is frustrated because this is the second attempt\nNext Best Action: Monitor until 48 hours without movement, then offer expedited replacement or 15% concession and escalate to Tier 2 Returns if needed
Stateless support handoff
text
Customer: Jordan Lee reported ORDER-7718 arrived with a cracked monitor stand.\nIssue: Damaged item; monitor stand is cracked on arrival.\nRequested Resolution: Customer wants a replacement shipped this week.\nOpen Question: Jordan asked whether the damaged item must be returned before replacement is sent.\nStatus: Damage claim captured and awaiting next-agent confirmation on replacement timing and return requirement.
State strategy recommendation
text
Recommendation: Stateless continuation\nReason: In a regulated support workflow, requiring names, order IDs, and refund context to be explicitly provided each turn improves controllability, auditability, and data-minimization, reducing the risk of unintended retention or cross-session leakage.
Hi Maya, I’m sorry your replacement order ORDER-8831 is delayed. I’m checking the latest replacement and carrier status now so I can confirm the best next step for you without making you wait longer than necessary.
Background manager summary
text
theme: Shipping delays dominate, especially West Coast distribution lane.\nrisk: Rising dissatisfaction from delays, replacements, and return exceptions.\nnext action: Escalate West Coast lane issues and review holiday return policy.
Compacted support context
json
{\n "feature": "Compaction",\n "how_to_apply": "Summarize older support turns into durable facts, open questions, policy constraints, and next actions before continuing the workflow.",\n "brightcart_example": {\n "durable_facts": [\n "Customer Maya Chen",\n "ORDER-8831",\n "replacement delayed",\n "carrier scan stale"\n ],\n "policy_constraints": [\n "Do not promise refund without eligibility",\n "Offer expedited replacement or 15% concession after 48-hour delay with approval"\n ],\n "next_action": "Check latest carrier scan and supervisor callback status."\n }\n}
example response_type \
0 Endpoint verification text
1 First raw HTTPS request text
2 SDK text generation text
3 Create and retrieve response text
4 Service tier and prompt cache request text
5 Structured ticket triage json
6 JSON support handoff json
7 Compact policy guidance text
8 Detailed policy guidance text
9 Order-status tool answer text
10 Parallel order lookup fallback answer text
11 Normalized support note text
12 Support transcript extraction text text
13 Stateful support handoff text
14 Stateless support handoff text
15 State strategy recommendation text
16 Prompt-cache token comparison table
17 Cached support-policy reply text
18 Background manager summary text
19 Compacted support context json
20 Endpoint responsiveness summary json
response
0
… (salida recortada)
example
response_type
response
0
Endpoint verification
text
ok
1
First raw HTTPS request
text
Empathy: I’m sorry, Maya — your replacement order ORDER-8831 is delayed because the carrier reported a temporary transit hold at the regional sorting facility.\nAction: We’re monitoring the shipment closely and will send you an updated delivery estimate within 24 hours; if there’s no movement by then, we’ll review the next replacement or refund options with you.
2
SDK text generation
text
Use the Responses API to build a BrightCart support assistant that can answer customer questions, summarize policies, and guide users through common workflows like order tracking, refunds, and account updates. Ground the assistant in BrightCart documentation and connect it to relevant backend tools or APIs so it can retrieve live order data, check account status, and provide accurate, context-aware support responses. Design the experience around clear system instructions, structured tool calling, and conversation state management so the assistant stays on-brand, reliable, and safe when handling customer issues.
3
Create and retrieve response
text
goal: Help support agents explain delayed replacement orders, set expectations, and suggest next steps.\ndata needed: Order ID, replacement order status, shipment/tracking events, delay reason, estimated ship/delivery date, customer contact history, inventory/backorder status, and applicable refund or reship policy.\nhuman-review rule: Escalate to a human if the delay exceeds policy thresholds, tracking is inconsistent or missing, the order appears lost, the customer is high-risk or highly upset, or any refund/reship exception is requested.
4
Service tier and prompt cache request
text
Latency benefit: Prompt caching lets the BrightCart support assistant reuse previously processed context, reducing response time for repeated or similar requests.\nConsistency benefit: Prompt caching helps the BrightCart support assistant return more uniform answers by reusing the same established prompt context across interactions.
5
Structured ticket triage
json
{\n "ticket_id": "TICKET-7429",\n "category": "delivery_delay",\n "priority": "urgent",\n "customer_sentiment": "frustrated and time-sensitive",\n "summary": "Customer Maya Chen reports that ORDER-8831 is a replacement shipment for a previously damaged standing desk. The replacement is now 2 days late, carrier tracking has not updated, and she needs the desk delivered before Monday. She is requesting a supervisor callback and wants to know refund options if the replacement cannot arrive in time.",\n "required_actions": [\n "Review ORDER-8831 shipment status and confirm last carrier scan/update.",\n "Contact carrier or open a trace/escalation for stalled tracking.",\n "Check expedited reshipment or alternative fulfillment options to meet the before-Monday deadline.",\n "Arrange supervisor callback per customer request.",\n "Review and communicate refund options, including refund\n...
6
JSON support handoff
json
{\n "customer_name": "Maya Chen",\n "order_id": "ORDER-8831",\n "issue_summary": "Customer is asking about a delayed replacement order. The carrier tracking scan is stale and has not updated.",\n "next_step": "Handoff to support to investigate the carrier delay, verify shipment status, and provide Maya Chen with an update or resolution.",\n "metrics_to_watch": [\n "tracking_scan_recency",\n "carrier_exception_status",\n "replacement_order_delivery_eta",\n "customer_follow_up_time"\n ]\n}
7
Compact policy guidance
text
BrightCart’s delayed-replacement policy lets customers keep using the original item until the replacement arrives, then return the defective product within the allowed return window.
8
Detailed policy guidance
text
1. BrightCart sends replacements after customers return the original item and warehouse receipt is confirmed.\n2. This delay prevents duplicate shipments, verifies eligibility, and reduces fraud or inventory errors.\n3. Agents should explain timelines clearly, offer return instructions, and reassure customers once receipt is logged.
9
Order-status tool answer
text
Status: ORDER-8831 is delayed; carrier shows no movement for 36 hours at the Denver sort center, with promised delivery on 2026-06-01.\nNext best action: Monitor until the 48-hour threshold; if no movement then, contact the customer and offer either an expedited replacement or a 15% concession with agent approval.
10
Parallel order lookup fallback answer
text
Order statuses: ORDER-8831 is delayed and ORDER-2044 was delivered yesterday.\nShipping problems: Maya has one active shipping problem, because only ORDER-8831 is currently delayed while ORDER-2044 is already delivered.
{"ticket_id":"TICKET-7429","customer":"Maya Chen","order_id":"ORDER-8831","product":"Standing desk replacement","issue":"Replacement for a damaged item is delayed and carrier scan has not moved","requested_resolution":"Supervisor callback and refund options","policy_options":"expedited replacement or 15% concession with agent approval after 48-hour delay"}
13
Stateful support handoff
text
Ticket ID: TICKET-4812\nOrder ID: ORDER-8831\nCustomer Name: Maya Chen\nIssue: Replacement standing desk shipment for damaged delivery has had no carrier movement for 36 hours; customer is frustrated because this is the second attempt\nNext Best Action: Monitor until 48 hours without movement, then offer expedited replacement or 15% concession and escalate to Tier 2 Returns if needed
14
Stateless support handoff
text
Customer: Jordan Lee reported ORDER-7718 arrived with a cracked monitor stand.\nIssue: Damaged item; monitor stand is cracked on arrival.\nRequested Resolution: Customer wants a replacement shipped this week.\nOpen Question: Jordan asked whether the damaged item must be returned before replacement is sent.\nStatus: Damage claim captured and awaiting next-agent confirmation on replacement timing and return requirement.
15
State strategy recommendation
text
Recommendation: Stateless continuation\nReason: In a regulated support workflow, requiring names, order IDs, and refund context to be explicitly provided each turn improves controllability, auditability, and data-minimization, reducing the risk of unintended retention or cross-session leakage.
Hi Maya, I’m sorry your replacement order ORDER-8831 is delayed. I’m checking the latest replacement and carrier status now so I can confirm the best next step for you without making you wait longer than necessary.
18
Background manager summary
text
theme: Shipping delays dominate, especially West Coast distribution lane.\nrisk: Rising dissatisfaction from delays, replacements, and return exceptions.\nnext action: Escalate West Coast lane issues and review holiday return policy.
19
Compacted support context
json
{\n "feature": "Compaction",\n "how_to_apply": "Summarize older support turns into durable facts, open questions, policy constraints, and next actions before continuing the workflow.",\n "brightcart_example": {\n "durable_facts": [\n "Customer Maya Chen",\n "ORDER-8831",\n "replacement delayed",\n "carrier scan stale"\n ],\n "policy_constraints": [\n "Do not promise refund without eligibility",\n "Offer expedited replacement or 15% concession after 48-hour delay with approval"\n ],\n "next_action": "Check latest carrier scan and supervisor callback status."\n }\n}
Lección del curso «OpenAI Cookbook» de OpenAI, publicado con licencia MIT. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a OpenAI. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios