Cuando implementas Llama para tu caso de uso, es una buena práctica tener Evals para este. Aunque lo ideal sería tener Evals anotados por humanos, este notebook muestra una estrategia sobre cómo abordar esto usando datos sintéticos. Sin embargo, los Evals generados aún requieren validación humana para asegurar que tu caso de uso en producción pueda depender de ellos.
El notebook también muestra cómo se pueden medir con precisión las alucinaciones sin usar la metodología LLM-As-A-Judge con Llama.
Idea general
Asumamos que tenemos un caso de uso para generar un informe de resumen basado en un contexto dado, lo cual es un caso de uso bastante común con LLM. Tanto el contexto como el informe tienen mucha información fáctica y queremos asegurarnos de que el informe generado no esté alucinando.
Dado que no es trivial encontrar un conjunto de datos de código abierto para esto, la idea es tomar datos tabulares sintéticos y luego usar Llama para generar una historia (contexto) para cada fila de los datos tabulares usando Prompt Engineering. Luego le pedimos a Llama que resuma el contexto generado como un informe en un formato específico usando Prompt Engineering. Finalmente, verificamos la precisión fáctica del informe generado usando Llama, convirtiendo esto en una tarea de QA usando los datos tabulares como la verdad fundamental.
Para generar datos sintéticos para este enfoque, usamos una herramienta de código abierto como Synthetic Data Vault
El flujo de trabajo general se muestra en el siguiente diagrama:
Instalación de Synthetic Data Vault
!pip install sdv
SDV tiene varios conjuntos de datos de una sola tabla. Elegimos el conjunto de datos student_placements para este notebook.
from sdv.datasets.demo import get_available_demos
get_available_demos(modality='single_table')
# Save the DataFrame to a CSV file
synthetic_data.to_csv('generated_data/tabular_data.csv')
Cargar datos tabulares sintéticos pregenerados
import pandas as pd
# Read the CSV file into a DataFrame
synthetic_data = pd.read_csv('generated_data/tabular_data.csv')
Generación de datos sintéticos con Llama-3.3-70B-Instruct
En esta sección, usamos Llama-3.3-70B-Instruct para crear una historia usando datos tabulares y luego generar un informe de resumen extractivo a partir del contexto generado.
Podrías intentar usar Llama-3.1-8B-Instruct, pero hemos visto mejores resultados con el modelo 70B para generar datos sintéticos.
Enfoque alternativo
En la sección siguiente, elegimos datos tabulares como la verdad fundamental y generamos todo el contexto y los informes a partir de la tabla. Otro enfoque es usar un par de ejemplos como few-shot prompting y usar Llama para generar el contexto y la historia a partir de esto, pidiéndole que varíe la información fáctica. Luego podemos usar Llama para crear los datos tabulares de la verdad fundamental.
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id: str = "meta-llama/Llama-3.3-70B-Instruct"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
/home/agunapal/anaconda3/envs/torchtune/lib/python3.10/site-packages/tqdm/auto.py:21: TqdmWarning: IProgress not found. Please update jupyter and ipywidgets. See https://ipywidgets.readthedocs.io/en/stable/user_install.html
from .autonotebook import tqdm as notebook_tqdm
Loading checkpoint shards: 100%|██████████| 30/30 [19:53<00:00, 39.79s/it]
# Prompt for generating context from synthetic tabular data
story_teller = lambda index: f"""You are expert Story Teller.
Look at the following data and tell a story in the form of a progress report.
This report should have a sub-section for the following:
- Academic Background
- Career Aspirations
- Salary Expectations
- Placement Status
- Course Details
- Story Behind the Numbers
Be creative and make up story with other statistics
{synthetic_data.loc[[synthetic_data.index[index]]]}
- DO NOT create another table.
- DO NOT ask any clarifying questions
- DO NOT justify your answers
- Make sure each column has a subheading in the report
- Each of the sections should have the respective tag.
- All currency is in tokens
- Answer within 800 tokens
Example:
<academic_background>
Student 17264 has a commendable academic record, which is evident from his second-year and high school percentage scores. His second-year percentage stands at 67%, while his high school percentage is an impressive 91%. He has also shown a keen interest in commerce, with a degree percentage of 58%. His academic background is a testament to his hard work and dedication to his studies.
</academic_background>
Answer:
"""
# Prompt for generating report from the generated context
report_creator = lambda context: f""" You are an expert report creator.
Look at the data in context:
{context}
and generate a shortened report with 1 line with the following subsections:
- student_id
- degree_type
- salary
- mba_spec
- duration
- employability_perc
IMPORTANT:
- DO NOT ask any clarifying questions
- DO NOT justify your answers
- DO NOT show the data
- DO NOT write any python code
- Each of the sections should have the respective tag and should be shown ONLY once
- Make sure to copy the mba_spec & degree_type as is
Example:
Summary Report:
<student_id>
Student ID is 17269
<student_id>
<salary>
Student has a realistic salary expectation of 27,000 tokens per month
<salary>
<degree_type>
Student has a degree in Sci&Tech
<degree_type>
<mba_spec>
Student has a specialization in Mkt&Fin
<mba_spec>
<duration>
Student has a degree duration of 4 years
<duration>
<employability_perc>
Student has a 95.0% employability percentage
<employability_perc>
Answer:
"""
Generar 12 ejemplos de datos sintéticos usando este bucle
¿Por qué 12?: Usaremos 2 ejemplos para few-shot prompting y los 10 restantes para Evals.
En la práctica, querrás que el número de puntos de datos sea mucho mayor para tu aplicación de producción.
import random
import json
for i in range(12):
formatted_prompt = story_teller(i)
input = tokenizer([formatted_prompt], return_tensors="pt").to("cuda")
# Generate context from tabular data
output = model.generate(**input, max_new_tokens=800, pad_token_id=0, temperature=0.8)
prompt_len = input["input_ids"].shape[-1]
context = tokenizer.decode(output[0][prompt_len:], skip_special_tokens=True)
formatted_prompt = report_creator(context)
input = tokenizer([formatted_prompt], return_tensors="pt").to("cuda")
# Generate report from generated report
output = model.generate(**input, max_new_tokens=120, pad_token_id=0,)
prompt_len = input["input_ids"].shape[-1]
report = tokenizer.decode(output[0][prompt_len:], skip_special_tokens=True)
# Create json output
result = {}
result["context"] = context
result["report"] = report
with open(f'generated_data/data_{i}.json', 'w') as f:
json.dump(result, f, indent=4)
Contexto e informe de ejemplo
Mediante inspección manual, vemos que Llama ha creado un contexto bien estructurado y el informe correspondiente. También vemos que toda la información fáctica es correcta.
import json
def read_json_file(file_path):
try:
with open(file_path, 'r') as file:
data = json.load(file)
return data
except FileNotFoundError:
print(f"File not found: {file_path}")
return None
except json.JSONDecodeError as e:
print(f"Invalid JSON: {e}")
return None
# Example usage:
file_path = 'generated_data/data_0.json'
data = read_json_file(file_path)
print("Context is -------------------------\n")
print(data["context"])
print("\nReport is -------------------------\n")
print(data["report"])
Context is -------------------------
### Progress Report for Student 3040587
#### <academic_background>
Student 3040587 has a strong academic foundation, with a high school percentage of 66.62% in Science. He also holds a degree in Science and Technology with a percentage of 75.76%. His second-year percentage is 75.01%, demonstrating his consistent academic performance.
#### <career_aspirations>
With a specialization in Marketing and Finance, Student 3040587 aspires to pursue a career in the finance sector, leveraging his skills in market analysis and financial planning. His career goal is to become a financial analyst, with a focus on investment banking.
#### <salary_expectations>
Student 3040587 expects a starting salary of 5000 tokens per annum, considering his one year of work experience and academic achievements. He is confident that his skills and knowledge will enable him to secure a job with a reputable company.
#### <placement_status>
Student 3040587 has been successfully placed, with an employability percentage of 85.98%. His placement is a testament to his hard work and dedication to his studies, as well as his relevant work experience.
#### <course_details>
Student 3040587 is currently pursuing an MBA with a specialization in Marketing and Finance, with a course duration of 3 years. He has completed one year of the course, with an MBA percentage of 58.37%.
#### <story_behind_the_numbers>
Behind the numbers, Student 3040587's story is one of perseverance and determination. Despite facing challenges in his academic journey, he has consistently worked hard to achieve his goals. His work experience has equipped him with the skills and knowledge required to succeed in the finance sector. With his strong academic background, career aspirations, and relevant work experience, Student 3040587 is poised to achieve great things in his future career.
### End of Report 3040587
Report is -------------------------
Summary Report:
<student_id>
Student 3040587
<student_id>
<salary>
Student has a realistic salary expectation of 5000 tokens per annum
<salary>
<degree_type>
Student has a degree in Science and Technology
<degree_type>
<mba_spec>
Student has a specialization in Marketing and Finance
<mba_spec>
<duration>
Student has a degree duration of 3 years
<duration>
<employability_perc>
Student has a 85.98% employability percentage
<employability_perc>
En este punto, idealmente necesitas que un humano revise los datos sintéticos que has generado y corrija cualquier error en el formato o la información fáctica, o que esté al tanto del número de errores en el conjunto de datos.
Medición de alucinaciones
El método habitual para medir alucinaciones utiliza la metodología LLM-As-Judge. Un ejemplo de métrica de alucinación es usar DeepEval.
Esto usaría un LLM potente como la verdad fundamental para medir las alucinaciones.
La sección siguiente muestra una forma de medir las alucinaciones utilizando los datos de la verdad fundamental que tenemos (datos tabulares). La metodología consiste en utilizar las etiquetas que hemos añadido en el informe y usar Llama para responder preguntas sencillas consultando las secciones correspondientes. Llama compara las respuestas con la verdad fundamental y genera una lista de valores booleanos. Esto se utiliza luego para medir la precisión de la información fáctica en el informe. Si tu informe tiene una estructura bien definida, usar QA para medir las alucinaciones puede ser muy efectivo y rentable.
# Use the first 2 data points for few shot prompting
file_path = 'generated_data_faulty/data_0.json'
example_data = read_json_file(file_path)
file_path = 'generated_data_faulty/data_1.json'
example_data_1 = read_json_file(file_path)
check_hallucinations = lambda data,index: f"""You are a Helpful Assistant.
Look at the section called Generated Report below & answer the following questions by only looking
at the corresponding sections in the report
- student_id: Question : What is the student id?
- degree_type : Question : What is the degree_type?
- salary: Question : What is the salary?
- mba_spec: Question : What is the mba_spec?
- duration: Question : What is the duration?
- employability_perc: Question : What is the employability percentage?
Generated Report:
{data["report"]}
Compare your answers with the ground truth and return either True or False within the tags <answer> & </answer>
Only if an answer is False, explain why in the format shown in the examples below
Ground Truth:
{synthetic_data.loc[[synthetic_data.index[index]]]}
Important Notes:
- Only check for the above mentioned questions
- Make sure each of the section is shown ONLY once
- DO NOT reason or explain your process
- DO NOT code this
- DO NOT explain why something is True
- Be lenient when checking decimal points. Ex: 4.0 is the same as 4
Example:
1)
With the following report:
{example_data["report"]}
and the ground truth:
{synthetic_data.loc[[synthetic_data.index[0]]]}
the following output is expected:
<answer>
student_id: [False, report shows 17263 and ground truth says 17264]
degree_type: [True, None]
salary: [False, report says 28000 and ground truth says 27000]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
2)
With the following report:
{example_data_1["report"]}
and the ground truth:
{synthetic_data.loc[[synthetic_data.index[1]]]}
the following output is expected:
<answer>
student_id: [True, None]
degree_type: [True, None]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Answer:
"""
def parse_output(output):
"""
Parse the output and return a list of bool values
"""
lines = output.strip().splitlines()
bool_values = []
for line in lines:
# Skip empty lines, lines with tags, or lines starting with '<'
if not line or line.startswith('<') or line.endswith('>'):
continue
parts = line.split(': ')
if len(parts) != 2:
raise ValueError(f"Invalid line format: {line}")
value_str, _ = parts[1].strip('[]').split(', ')
if value_str == 'True':
bool_values.append(True)
elif value_str == 'False':
bool_values.append(False)
else:
raise ValueError(f"Invalid bool value: {value_str}")
if parts[0] == 'employability_perc':
break
return bool_values
from sklearn.metrics import accuracy_score
y_pred = []
y_true = [True]*60
for i in range(2,12):
fname = f'generated_data/data_{i}.json'
print(f"\nChecking accuracy of generated report in {fname}\n")
data = read_json_file(fname)
formatted_prompt = check_hallucinations(data, i)
input = tokenizer([formatted_prompt], return_tensors="pt").to("cuda")
output = model.generate(**input, max_new_tokens=120, pad_token_id=0, do_sample=False, top_p=None, temperature=None)
prompt_len = input["input_ids"].shape[-1]
results = tokenizer.decode(output[0][prompt_len:], skip_special_tokens=True)
print(results)
y_pred.extend(parse_output(results))
accuracy = accuracy_score(y_true, y_pred)
print(f"\nAccuracy of factual information generation is : {accuracy:.4f}")
Checking accuracy of generated report in generated_data/data_2.json
<answer>
student_id: [True, None]
degree_type: [True, None]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Checking accuracy of generated report in generated_data/data_3.json
<answer>
student_id: [True, None]
degree_type: [True, None]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Checking accuracy of generated report in generated_data/data_4.json
<answer>
student_id: [True, None]
degree_type: [True, None]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Checking accuracy of generated report in generated_data/data_5.json
<answer>
student_id: [True, None]
degree_type: [True, None]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Checking accuracy of generated report in generated_data/data_6.json
<answer>
student_id: [False, report shows Student ID is not mentioned in the data and ground truth says 6180804]
degree_type: [True, None]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Checking accuracy of generated report in generated_data/data_7.json
<answer>
student_id: [True, None]
degree_type: [False, report says Mkt&Fin and ground truth says Sci&Tech]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Checking accuracy of generated report in generated_data/data_8.json
<answer>
student_id: [True, None]
degree_type: [True, None]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Checking accuracy of generated report in generated_data/data_9.json
<answer>
student_id: [True, None]
degree_type: [True, None]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Checking accuracy of generated report in generated_data/data_10.json
<answer>
student_id: [True, None]
degree_type: [True, None]
salary: [True, None]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Checking accuracy of generated report in generated_data/data_11.json
<answer>
student_id: [True, None]
degree_type: [False, report says Commerce and ground truth says Comm&Mgmt]
salary: [False, report says 500 and ground truth says NaN]
mba_spec: [True, None]
duration: [True, None]
employability_perc: [True, None]
</answer>
Accuracy of factual information generation is : 0.9333
Conclusión y próximos pasos
Crear Evals para la sumarización es importante.
Llama se puede usar para crear Evals dados algunos ejemplos de la verdad fundamental.
Usar QA simple para medir alucinaciones puede ser una estrategia efectiva para confiar en que la información fáctica importante se está verificando.
Lección del curso «Llama Cookbook (use cases)» de Meta, publicado con licencia MIT. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Meta. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios