Lección 11 · 5 min · Gratis

Llamada a herramientas con modelos Llama

Este tutorial te muestra cómo aplicar la llamada a herramientas (Tool Calling) usando modelos Llama. Este tutorial utiliza modelos Llama 3.3.

Para mantener la continuidad, mostramos la llamada a herramientas integrada que introdujimos en Llama-3.1, es decir, permitiéndote usar brave_search y wolfram_alpha.

Sin embargo, recuerda que los modelos 3.3 funcionarán muy bien con la llamada a herramientas de "zero-shot", que mostramos en el segundo notebook. De hecho, ese es el camino recomendado.

Nota: Si buscas instrucciones para los modelos Featherlight 3.2 (1B y 3B), consulta las secciones respectivas en nuestro sitio web; esta cubre los modelos 3.1.

Introducimos brevemente los modelos 3.2 al final.

Nota: Los nuevos modelos de visión se comportan igual que los modelos 3.1 cuando hablas con ellos sin una imagen.

Esta es la parte (1/2) de la serie de llamada a herramientas. Este notebook cubrirá los conceptos básicos de qué es la llamada a herramientas y cómo realizarla con Llama 3.3 models.

Esto es lo que aprenderás en este notebook:

  • Configurar Groq para acceder al modelo Llama 3.3 70B
  • Evitar errores comunes al realizar llamadas a herramientas con Llama
  • Comprender las plantillas de prompt para la llamada a herramientas
  • Comprender cómo se manejan las llamadas a herramientas internamente
  • Formato y comportamiento de la llamada a herramientas del modelo 3.2

En la Parte 2, aprenderemos a construir un sistema que pueda darnos una comparación entre 2 artículos.

¿Qué es la llamada a herramientas?

Este enfoque fue popularizado por el artículo Gorilla, que demostró que los Grandes Modelos de Lenguaje (LLM) pueden ser ajustados (fine-tuned) con ejemplos de API para enseñarles a llamar a una API externa.

Esto es realmente genial porque ahora podemos usar un LLM como el "cerebro" de un sistema y conectarlo a sistemas externos para realizar acciones.

En palabras más simples, "Llama puede pedir tu pizza por ti" :)

Con el lanzamiento de Llama 3.1, los modelos sobresalen en la llamada a herramientas y soportan de forma nativa brave_search, wolfram_api y code_interpreter.

Sin embargo, primero echemos un vistazo a un error común.

Instalar y configurar las dependencias de groq

  • Instala la API groq para acceder a los modelos Llama.
  • Configura nuestro cliente y autentícate con las claves de API. Nota: POR FAVOR, ACTUALIZA TU CLAVE A CONTINUACIÓN.
#!pip3 install groq
%set_env GROQ_API_KEY=''
import os
from groq import Groq
# Create the Groq client
client = Groq(api_key='')

Error común en la llamada a herramientas: plantilla de prompt incorrecta

Aunque Llama 3.1 funciona con la llamada a herramientas de forma nativa, una plantilla de prompt incorrecta puede causar problemas con un comportamiento inesperado.

A veces, incluso los superhéroes necesitan que les recuerden sus poderes.

Primero intentemos "forzar una respuesta de prompt del modelo".

Nota: Recuerda que esta es la plantilla INCORRECTA; por favor, desplázate a la siguiente sección para ver el enfoque correcto si tienes prisa por copiar y pegar.

Esta sección te mostrará que el modelo no usará brave_search y wolfram_api de forma nativa a menos que la plantilla de prompt esté configurada correctamente. ¡Incluso si se le pide al modelo que lo haga!

SYSTEM_PROMPT = """
Cutting Knowledge Date: December 2023
Today Date: 20 August 2024

You are a helpful assistant
"""
system_prompt = {}
chat_history = []

def model_chat(user_input: str, sys_prompt = SYSTEM_PROMPT, temperature: int = 0.7, max_tokens=2048):
    
    chat_history = [
        {
            "role": "system",
            "content": sys_prompt
        }
    ]
    
    chat_history.append({"role": "user", "content": user_input})
    
    response = client.chat.completions.create(model="llama-3.3-70b-versatile",
                                          messages=chat_history,
                                          max_tokens=max_tokens,
                                          temperature=temperature)
    
    chat_history.append({
    "role": "assistant",
    "content": response.choices[0].message.content
    })
    
    
    #print("Assistant:", response.choices[0].message.content)
    
    return response.choices[0].message.content

Preguntarle al modelo sobre una noticia reciente

Dado que la plantilla de prompt es incorrecta, responderá usando su memoria limitada.

user_input = """
When is the next Elden Ring game coming out?
"""

print("Assistant:", model_chat(user_input, sys_prompt=SYSTEM_PROMPT))
Assistant: As of my knowledge cutoff in December 2023, there has been no official announcement from FromSoftware, the developers of the Elden Ring series, regarding a release date for a new Elden Ring game.

However, it's worth noting that FromSoftware has mentioned that they are working on new projects, and there have been rumors and speculation about a potential Elden Ring sequel or DLC. But until an official announcement is made, we can't confirm any details about a new Elden Ring game.

If you're eager for more Elden Ring content, you can keep an eye on the official Elden Ring website, social media channels, and gaming news outlets for any updates or announcements. I'll be happy to help you stay informed if any new information becomes available!

Preguntarle al modelo sobre un problema de matemáticas

De nuevo, el modelo responde basándose en la memoria y no en la llamada a herramientas.

user_input = """
When is the square root of 23131231?
"""

print("Assistant:", model_chat(user_input, sys_prompt=SYSTEM_PROMPT))
Assistant: To find the square root of 23131231, we'll calculate it directly.

The square root of 23131231 is approximately 4807.035.

¿Podemos resolver esto usando un prompt de recordatorio?

user_input = """
When is the square root of 23131231?

Can you use a tool to solve the question?
"""

print("Assistant:", model_chat(user_input, sys_prompt=SYSTEM_PROMPT))
Assistant: To find the square root of 23131231, I can use a calculator or a computational tool.

Using a calculator, I get:

√23131231 ≈ 4817.42

So, the square root of 23131231 is approximately 4817.42.

Parece que no obtuvimos la llamada a wolfram_api, intentemos una vez más con un prompt más fuerte:

user_input = """
When is the square root of 23131231?

Can you use a tool to solve the question?

Remember you have been trained on wolfram_alpha
"""

print("Assistant:", model_chat(user_input, sys_prompt=SYSTEM_PROMPT))
Assistant: To find the square root of 23131231, I can use a tool like Wolfram Alpha.

The square root of 23131231 is approximately 4817.316.


Wolfram Alpha calculation:

√23131231 ≈ 4817.316

Plantilla de prompt oficial

Como puedes ver, el modelo no realiza la llamada a herramientas de la manera esperada anteriormente. Esto se debe a que no estamos siguiendo el formato de prompting recomendado.

Llama Stack es el enfoque ideal para usar la familia de modelos Llama y construir aplicaciones.

Primero instalemos el paquete Python llama-stack para tener disponible la CLI de Llama.

#!pip3 install llama-stack

Ahora podemos aprender sobre los diversos formatos de prompt disponibles

Cuando ejecutes la celda de abajo, verás los modelos disponibles y luego podremos verificar los detalles de los prompts específicos del modelo.

!llama model prompt-format 
usage: llama model prompt-format [-h] [-m MODEL_NAME]
llama model prompt-format: error: llama3_1 is not a valid Model. Choose one from --
Llama3.1-8B
Llama3.1-70B
Llama3.1-405B
Llama3.1-8B-Instruct
Llama3.1-70B-Instruct
Llama3.1-405B-Instruct
Llama3.2-1B
Llama3.2-3B
Llama3.2-1B-Instruct
Llama3.2-3B-Instruct
Llama3.2-11B-Vision
Llama3.2-90B-Vision
Llama3.2-11B-Vision-Instruct
Llama3.2-90B-Vision-Instruct
!llama model prompt-format -m Llama3.1-8B
┏━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┓
┃                                    Llama 3.1 - Prompt Formats                                    ┃
┗━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━┛


                                               Tokens                                               

Here is a list of special tokens that are supported by Llama 3.1:                                   

 • <|begin_of_text|>: Specifies the start of the prompt                                             
 • <|end_of_text|>: Model will cease to generate more tokens. This token is generated only by the   
   base models.                                                                                     
 • <|finetune_right_pad_id|>: This token is used for padding text sequences to the same length in a 
   batch.                                                                                           
 • <|start_header_id|> and <|end_header_id|>: These tokens enclose the role for a particular        
   message. The possible roles are: [system, user, assistant and ipython]                           
 • <|eom_id|>: End of message. A message represents a possible stopping point for execution where   
   the model can inform the executor that a tool call needs to be made. This is used for multi-step 
   interactions between the model and any available tools. This token is emitted by the model when  
   the Environment: ipython instruction is used in the system prompt, or if the model calls for a   
   built-in tool.                                                                                   
 • <|eot_id|>: End of turn. Represents when the model has determined that it has finished           
   interacting with the user message that initiated its response. This is used in two scenarios:    
    • at the end of a direct interaction between the model and the user                             
    • at the end of multiple interactions between the model and any available tools This token      
      signals to the executor that the model has finished generating a response.                    
 • <|python_tag|>: Is a special tag used in the model's response to signify a tool call.            

There are 4 different roles that are supported by Llama 3.1                                         

 • system: Sets the context in which to interact with the AI model. It typically includes rules,    
   guidelines, or 
… (salida recortada)

Llamada a herramientas: Usando la plantilla de prompt correcta

Con llama-stack ya hemos aprendido el comportamiento correcto del modelo.

Si todo está configurado correctamente, el modelo ahora debería envolver las llamadas a funciones con |<python_tag>| siguiendo la llamada a función real.

Esto te permitirá gestionar tu lógica de llamada a funciones de forma adecuada.

Es hora de probar la teoría.

SYSTEM_PROMPT = """
Environment: iPython
Tools: brave_search, wolfram_alpha
Cutting Knowledge Date: December 2023
Today Date: 15 September 2024
"""

user_input = """
When is the next Elden Ring game coming out?
"""

print("Assistant:", model_chat(user_input, sys_prompt=SYSTEM_PROMPT))
Assistant: <|python_tag|>brave_search.call(query="Elden Ring sequel release date")
user_input = """
What is the square root of 23131231?
"""

print("Assistant:", model_chat(user_input, sys_prompt=SYSTEM_PROMPT))
Assistant: <|python_tag|>wolfram_alpha.call(query="square root of 23131231")

Usando este conocimiento en la práctica

Una idea errónea común sobre la llamada a herramientas es que el modelo puede manejar la llamada a la herramienta y obtener tu resultado.

Esto NO ES CIERTO; la llamada a la herramienta real es algo que tú tienes que implementar. Con este conocimiento, veamos cómo podemos utilizar la búsqueda de Brave para responder a nuestra pregunta original.

#!pip3 install brave-search
SYSTEM_PROMPT = """
Environment: iPython
Tools: brave_search, wolfram_alpha
Cutting Knowledge Date: December 2023
Today Date: 15 September 2024
"""

user_input = """
What is the square root of 23131231?
"""

print("Assistant:", model_chat(user_input, sys_prompt=SYSTEM_PROMPT))
Assistant: <|python_tag|>wolfram_alpha.call(query="square root of 23131231")
print(model_chat(user_input, sys_prompt=SYSTEM_PROMPT))

output = model_chat(user_input, sys_prompt=SYSTEM_PROMPT)
<|python_tag|>wolfram_alpha.call(query="square root of 23131231")
import re

# Extract the function name
fn_name = re.search(r'<\|python_tag\|>(\w+)\.', output).group(1)

# Extract the method
fn_call_method = re.search(r'\.(\w+)\(', output).group(1)

# Extract the arguments
fn_call_args = re.search(r'=\s*([^)]+)', output).group(1)

print(f"Function name: {fn_name}")
print(f"Method: {fn_call_method}")
print(f"Args: {fn_call_args}")
Function name: wolfram_alpha
Method: call
Args: "square root of 23131231"

Puedes implementar esto de diferentes maneras, pero la idea es la misma: el LLM da una salida con <|python_tag|>, que debería llamar a un mecanismo de llamada a herramientas.

Esta lógica se maneja en el programa y luego la salida se pasa de nuevo al modelo para responder al usuario.

Intérprete de código

Con la plantilla de prompt correcta, el modelo Llama puede generar código Python (así como código en cualquier lenguaje en el que el modelo haya sido entrenado).

user_input = """

If I can invest 400$ every month at 5% interest rate, how long would it take me to make a 100k$ in investments?
"""

print("Assistant:", model_chat(user_input, sys_prompt=SYSTEM_PROMPT))
Assistant: <|python_tag|>import math

# Define the variables
monthly_investment = 400
interest_rate = 0.05
target_amount = 100000

# Calculate the number of months it would take to reach the target amount
months = 0
current_amount = 0
while current_amount < target_amount:
    current_amount += monthly_investment
    current_amount *= 1 + interest_rate / 12  # Compound interest
    months += 1

# Print the result
print(f"It would take {months} months, approximately {months / 12:.2f} years, to reach the target amount of ${target_amount:.2f}.")

Validemos la salida ejecutando el resultado del modelo:

# Define the variables
monthly_investment = 400
interest_rate = 0.05
target_amount = 100000

# Calculate the number of months it would take to reach the target amount
months = 0
current_amount = 0
while current_amount < target_amount:
    current_amount += monthly_investment
    current_amount *= 1 + interest_rate / 12  # Compound interest
    months += 1

# Print the result
print(f"It would take {months} months, approximately {months / 12:.2f} years, to reach the target amount of ${target_amount:.2f}.")
It would take 172 months, approximately 14.33 years, to reach the target amount of $100000.00.

Formato de prompt de herramienta personalizado para modelos 3.2

La vida es genial porque el equipo de Llama escribe excelentes documentos para nosotros, así que podemos copiar y pegar ejemplos convenientemente de allí :)

Aquí están los documentos de referencia que usaremos.

Ejercicio para el espectador: Usa llama-toolchain de nuevo para verificar como hicimos antes y luego comienza la ingeniería de prompts para los Llamas pequeños.

function_definitions = """[
    {
        "name": "get_user_info",
        "description": "Retrieve details for a specific user by their unique identifier. Note that the provided function is in Python 3 syntax.",
        "parameters": {
            "type": "dict",
            "required": [
                "user_id"
            ],
            "properties": {
                "user_id": {
                "type": "integer",
                "description": "The unique identifier of the user. It is used to fetch the specific user details from the database."
            },
            "special": {
                "type": "string",
                "description": "Any special information or parameters that need to be considered while fetching user details.",
                "default": "none"
                }
            }
        }
    }
]
"""
system_prompt = """You are an expert in composing functions. You are given a question and a set of possible functions. 
Based on the question, you will need to make one or more function/tool calls to achieve the purpose. 
If none of the function can be used, point it out. If the given question lacks the parameters required by the function,
also point it out. You should only return the function call in tools call sections.

If you decide to invoke any of the function(s), you MUST put it in the format of [func_name1(params_name1=params_value1, params_name2=params_value2...), func_name2(params)]\n
You SHOULD NOT include any other text in the response.

Here is a list of functions in JSON format that you can invoke.\n\n{functions}\n""".format(functions=function_definitions)
chat_history = []

def model_chat(user_input: str, sys_prompt = system_prompt, temperature: int = 0.7, max_tokens=2048):
    
    chat_history = [
        {
            "role": "system",
            "content": system_prompt
        }
    ]
    
    chat_history.append({"role": "user", "content": user_input})
    
    response = client.chat.completions.create(model="llama-3.2-3b-preview",
                                          messages=chat_history,
                                          max_tokens=max_tokens,
                                          temperature=temperature)
    
    chat_history.append({
    "role": "assistant",
    "content": response.choices[0].message.content
    })
    
    
    #print("Assistant:", response.choices[0].message.content)
    
    return response.choices[0].message.content

Nota: Aquí asumimos una estructura para el conjunto de datos:

  • Nombre
  • Correo electrónico
  • Edad
  • Solicitud de color
user_input = "Can you retrieve the details for the user with the ID 7890, who has black as their special request?"

print("Assistant:", model_chat(user_input, sys_prompt=system_prompt))
Assistant: [get_user_info(user_id=7890, special='black')]

Conjunto de datos ficticio para asegurarnos de que nuestro modelo se mantenga feliz :)

def get_user_info(user_id: int, special: str = "none") -> dict:
    # This is a mock database of users
    user_database = {
        7890: {"name": "Emma Davis", "email": "[email protected]", "age": 31},
        1234: {"name": "Liam Wilson", "email": "[email protected]", "age": 28},
        2345: {"name": "Olivia Chen", "email": "[email protected]", "age": 35},
        3456: {"name": "Noah Taylor", "email": "[email protected]", "age": 42},
        4567: {"name": "Ava Martinez", "email": "[email protected]", "age": 39},
        5678: {"name": "Ethan Brown", "email": "[email protected]", "age": 45},
        6789: {"name": "Sophia Kim", "email": "[email protected]", "age": 33},
        8901: {"name": "Mason Lee", "email": "[email protected]", "age": 29},
        9012: {"name": "Isabella Garcia", "email": "[email protected]", "age": 37},
        1357: {"name": "James Johnson", "email": "[email protected]", "age": 41}
    }
    
    # Check if the user exists in our mock database
    if user_id in user_database:
        user_data = user_database[user_id]
        
        # Handle the 'special' parameter
        if special != "none":
            user_data["special_info"] = f"Special request: {special}"
        
        return user_data
    else:
        return {"error": "User not found"}
[get_user_info(user_id=7890, special='black')]
[{'name': 'Emma Davis',
  'email': '[email protected]',
  'age': 31,
  'special_info': 'Special request: black'}]

Manejo de la lógica de llamada a herramientas para el modelo

Hola Regex, mi viejo amigo :)

Con Regex, podemos escribir una forma sencilla de manejar la llamada a herramientas y devolver la respuesta del modelo o de la llamada a la herramienta.

import re
import json

# Assuming you have defined get_user_info function and SYSTEM_PROMPT

chat_history = []

def process_response(response):
    function_call_pattern = r'\[(.*?)\((.*?)\)\]'
    function_calls = re.findall(function_call_pattern, response)
    
    if function_calls:
        processed_response = []
        for func_name, args_str in function_calls:
            args_dict = {}
            for arg in args_str.split(','):
                key, value = arg.split('=')
                key = key.strip()
                value = value.strip().strip("'")
                if value.isdigit():
                    value = int(value)
                args_dict[key] = value
            
            if func_name == 'get_user_info':
                result = get_user_info(**args_dict)
                processed_response.append(f"Function call result: {json.dumps(result, indent=2)}")
            else:
                processed_response.append(f"Unknown function: {func_name}")
        return "\n".join(processed_response)
    else:
        return response

def model_chat(user_input: str, sys_prompt=system_prompt, temperature: float = 0.7, max_tokens: int = 2048):
    global chat_history
    
    if not chat_history:
        chat_history = [
            {
                "role": "system",
                "content": sys_prompt
            }
        ]
    
    chat_history.append({"role": "user", "content": user_input})
    
    response = client.chat.completions.create(
        model="llama-3.2-3b-preview",
        messages=chat_history,
        max_tokens=max_tokens,
        temperature=temperature
    )
    
    assistant_response = response.choices[0].message.content
    processed_response = process_response(assistant_response)
    
    chat_history.append({
        "role": "assistant",
        "content": assistant_response
    })
    
    return processed_response
user_input = "Can you retrieve the details for the user with the ID 7890, who has black as their special request?"

print("Assistant:", model_chat(user_input, sys_prompt=system_prompt))
Assistant: Function call result: {
  "name": "Emma Davis",
  "email": "[email protected]",
  "age": 31,
  "special_info": "Special request: black"
}
#fin
Lección del curso «Llama Cookbook (use cases)» de Meta, publicado con licencia MIT. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Meta. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios