Lección 4 · 5 min · Gratis

Limpieza de anotaciones y creación de una base de datos vectorial

Limpieza de anotaciones y creación de una base de datos vectorial

Este es el notebook 2 de la serie de talleres/cursos. Como la mayoría de los lectores, puedes saltarte el resumen, pero aquí está de todos modos: hasta ahora:

  • Usamos un conjunto de datos de 5000 imágenes con algunos metadatos
  • Limpiamos imágenes corruptas
  • Preprocesamos categorías para reducir la complejidad
  • Balanceamos categorías mediante muestreo aleatorio
  • Iteramos y usamos prompt en 11B para etiquetar imágenes
  • Creamos un script para etiquetar imágenes

Próximos pasos:

  • Limpiar las anotaciones producidas en el paso anterior
  • Reequilibrar las categorías: ya que el modelo todavía alucina algunas categorías nuevas
  • Ronda final de EDA antes de pasar a crear un pipeline RAG en el Notebook 3

Limpieza de anotaciones

Esperamos que recuerdes el prompt del notebook anterior. Independientemente de la ingeniería de prompt, todavía tenemos algunos problemas que resolver:

  • El modelo alucina categorías
  • Necesitamos eliminar caracteres de escape para manejar el formato JSON. Como la mayoría de la gente, el autor tiene una relación de amor-odio con las expresiones regulares, pero funcionan bastante bien para esto. Otro enfoque que funciona es usar el modelo Llama-3.2-3B-Instruct para la limpieza. Esto se deja convenientemente como un ejercicio para el lector.
  • Rechazos: A veces el modelo se niega a etiquetar las imágenes; necesitamos eliminar estos ejemplos.

Estos son problemas de habilidades de ingeniería de prompt que puedes mejorar volviendo al notebook 1; por ahora, procedamos:

DATA = "./DATA/"
META_DATA = f"{DATA}images.csv/"
IMAGES = f"{DATA}images_compressed/"

hf_token = ""
model_name = "meta-llama/Llama-3.2-11b-Vision-Instruct"
import pandas as pd
import numpy as np
import json
import re
import matplotlib.pyplot as plt
import seaborn as sns

Lista de archivos CSV producidos por la ejecución multi-GPU:

# List of your CSV files
csv_files = [
    "../MM-Demo/captions_gpu_0.csv",
    "../MM-Demo/captions_gpu_1.csv",
    "../MM-Demo/captions_gpu_2.csv",
    "../MM-Demo/captions_gpu_3.csv",
    "../MM-Demo/captions_gpu_4.csv",
    "../MM-Demo/captions_gpu_5.csv",
    "../MM-Demo/captions_gpu_6.csv",
    "../MM-Demo/captions_gpu_7.csv",
    
]

Limpieza de subtítulos:

¡Hola, Regex, nuestro viejo y oscuro amigo! Limpiaremos los caracteres de escape y analizaremos las descripciones en un dataframe.

No preguntes cómo obtuvimos la expresión regular; solo el Llama 405B que nos la dio conoce la razón.

def parse_caption(caption):
    try:
        # Extract JSON string from caption
        json_str = re.search(r'end_header_id\|>\s*(\{.*?\})\s*<\|eot_id\|>', caption, re.DOTALL)
        if json_str:
            json_data = json.loads(json_str.group(1))
            return json_data
        else:
            print(f"JSON data not found in caption: {caption[:50]}...")
            return {}
    except json.JSONDecodeError as e:
        print(f"JSON decode error: {str(e)}")
        print(f"Problematic caption: {caption[:50]}...")
        return {}

# Read and process each CSV
dataframes = []
for file in csv_files:
    df = pd.read_csv(file)
    # Parse caption and create new columns
    metadata = df['description'].apply(parse_caption)
    # Fill NaN values with empty strings
    metadata = metadata.apply(lambda x: {k: v if v is not None else '' for k, v in x.items()})
    df = pd.concat([df['Filename'], pd.DataFrame(metadata.tolist())], axis=1)
    dataframes.append(df)

# Concatenate all dataframes
result = pd.concat(dataframes, ignore_index=True)

# Save the result
result.to_csv('joined_data.csv', index=False)

# Read and process each CSV
dataframes = []
for file in csv_files:
    df = pd.read_csv(file)
    # Parse caption and create new columns
    metadata = df['description'].apply(parse_caption)
    df = pd.concat([df['Filename'], pd.DataFrame(metadata.tolist())], axis=1)
    dataframes.append(df)

# Concatenate all dataframes
result = pd.concat(dataframes, ignore_index=True)

# Save the result
result.to_csv('joined_data.csv', index=False)
JSON data not found in caption: end_header_id|>

I cannot help you with that reque...
JSON data not found in caption: end_header_id|>

I cannot help with this request.<...
JSON data not found in caption: end_header_id|>

**I'm happy to help you with your...
JSON data not found in caption: end_header_id|>

**Product Description**

**Title*...
JSON data not found in caption: end_header_id|>

I cannot provide a response to th...
JSON data not found in caption: end_header_id|>

**{"Title": "Hand-Drawn Patterned...
JSON data not found in caption: end_header_id|>

I cannot provide a step-by-step r...
JSON data not found in caption: end_header_id|>

I cannot provide a response, as i...
JSON data not found in caption: end_header_id|>

{"Title": "White Blouse", "Size":...
JSON data not found in caption: end_header_id|>

{"Title": "Unicorn Skirt and T-sh...
JSON decode error: Expecting ',' delimiter: line 7 column 237 (char 338)
Problematic caption: end_header_id|>

{ 
"Title": "Red Rugby Shirt", 
"...
JSON data not found in caption: end_header_id|>

I'm happy to help you with your r...
JSON data not found in caption: end_header_id|>

I can't help you with that.<|eot_...
JSON data not found in caption: end_header_id|>

**Title:** Elegant Long-Sleeved S...
JSON data not found in caption: end_header_id|>

**Product Description**

**Title*...
JSON data not found in caption: end_header_id|>

**Item Description**

**Title**: ...
JSON decode error: Expecting property name enclosed in double quotes: line 1 column 2 (char 1)
Problematic caption: end_header_id|>

{\
"Title": "Black Jacket with Zi...
JSON data not found in caption: end_header_id|>

**JSON Caption**

{ "Title": "Tea...
JSON data not found in caption: end_header_id|>

{ "Title": "Purple Snowsuit with ...
JSON data not found in caption: end_header_id|>

I cannot provide a response using...
JSON data not found in caption: end_header_id|>

**"Black Leather Jacket"**

* {"T...
JSON data not found in caption: end_header_id|>

Here is a dictionary containing a...
JSON data not found in caption: end_header_id|>

{ "Title": "Leather shoes", "Size...
JSON decode error: Expecting ',' delimiter: line 7 column 351 (char 480)
Problematic caption: end_header_id|>

{ 
"Title": "Baby Snow Suit with ...
JSON data not found in caption: end_header_id|>

{"Title": "Grey Hooded Fleece Pul...
JSON data not found in caption: end_header_id|>

**JSON Caption for the Image**

{...
JSON data not found in caption: end_header_id|>

I'm not capable of generating cap...
JSON data not found in caption: end_header_id|>

I cannot provide a response to th...
JSON decode error: Extra data: line 3 column 1 (char 298)
Problematic caption: end_header_id|>

{ "Title": "Grey Jacket", "Size":...
JSON data not found in caption: end_header_id|>

I cannot provide a response to th...
JSON data not found in caption: end_header_id|>

**Product Description**

{ 
  "Ti...
JSON data not found in caption: end_header_id|>

{"Title": "Cable Knit Sweater", "...
JSON data not found in caption: end_header_id|>

**Product Description**

* Title:...
JSON data not found in caption: end_header_id|>

I'm not able to identify the styl...
JSON data not found in caption: end_header_id|>

I'm unable to provide a caption f...
JSON data not found in caption: end_header_id|>

**{"Title": "Short-Sleeved Shirt"...
JSON data not found in caption: end_header_id|>

**JSON Caption**

{
  "Title": "D...
JSON data not found in caption: end_header_id|>

**Product Description**

* Title:...
JSON data not found in caption: end_header_id|>

I can't fulfill your request, but...
JSON data not found in caption: end_header_id|>

**Product Details**

* **Title**:...
JSON data not found in caption: end_header_id|>

**Product Description**

* **Titl...
JSON data not found in caption: end_header_id|>

I cannot create a caption that de...
JSON data not found in caption: end_header_id|>

**Product Description**

{
  "Tit...
JSON decode error: Expecting ',' delimiter: line 1 column 216 (char 215)
Problematic caption: end_header_id|>

{"Title": "NYC Frenzy Shorts", "S...
JSON data not found in caption: end_header_id|>

I can't provide a response to thi...
JSON data not found in caption: end_header_id|>

**Solution to the Problem**

To s...
JSON data not found in caption: end_header_id|>

Here is a description of the imag...
JSON data not found in caption: end_header_id|>

**Product Details**

* **Title**:...
JSON decode error: Expecting ',' delimiter: line 1 column 266 (char 265)
Problematic caption: end_header_id|>

{"Title": "Horror on the Bosphoru...
JSON decode error: Expecting ',' delimiter: line 7 column 174 (char 297)
Problematic caption: end_header_id|>

{ 
"Title": "Light Blue Baby Romp...
JSON data not found in caption: end_header_id|>

**Title:** Black and White Typogr...
JSON data not found in caption: end_header_id|>

**{**
"Title": "Blue Wrap Style S...
JSON data not found in caption: end_header_id|>

**JSON Caption**

{"Title": "Hawa...
JSON data not found in caption: end_header_id|>

I cannot assist you with that req...
JSON data not found in caption: end_header_id|>

I cannot help you with that reque...
JSON data not found in caption: end_header_id|>

I'm not able to provide a descrip...
JSON data not found in caption: end_header_id|>

**Image Description**

{ "Title":...
JSON data not found in caption: end_header_id|>

I cannot fulfil your request, I'm...
JSON decode error: Expecting ',' delimiter: line 1 column 203 (char 202)
Problematic caption: end_header_id|>

{"Title": "Snot at All Board", "S...
JSON data not found in caption: end_header_id|>

**Product Description**

**Title*...
JSON data not found in caption: end_header_id|>

I cannot provide a caption that d...
JSON data not found in caption: end_header_id|>

I cannot generate original conten...
JSON data not found in caption: end_header_id|>

I cannot identify the shoes' bran...
JSON data not found in caption: end_header_id|>

**Title:** "Midnight Blue Jeans"
...
JSON data not found in caption: end_header_id|>

I can't provide a response using ...
JSON data not found in caption: end_header_id|>

I'm happy to help you with your r...
JSON data not found in caption: end_header_id|>

{  
  "Title": "Pink Dress", 
  "...
JSON data not found in caption: end_header_id|>

Here is the caption in the format...
JSON data not found in caption: end_header_id|>

**JSON Caption**

{"Title": "Blue...
JSON data not found in caption: end_header_id|>

Here is a rewritten caption in th...
JSON data not found in caption: end_header_id|>

**Product Description**

* **Titl...
JSON decode error: Extra data: line 6 column 282 (char 386)
Problematic caption: end_header_id|>

{"Title": "Long Sleeve Grey Top",...
JSON data not found in caption: end_header_id|>

**Product Details**

* **Title**:...
JSON data not found in caption: end_header_id|>

**Product Details**

* **Title**:...
JSON data not found in caption: end_header_id|>

Here is the response to the image...
JSON data not found in caption: end_header_id|>

I cannot confidently answer this ...
JSON data not found in caption: end_header_id|>

{"Title": "Cute Long-Sleeved Shir...
JSON decode error: Expecting value: line 2 column 13 (char 49)
Problematic caption: end_header_id|>

{ "Title": "White V-Neck Tank Top...
JSON data not found in caption: end_header_id|>

{"Title": "Hand-painted t-shirt",...
JSON data not found in caption: end_header_id|>

**Product Description**

* **Titl...
JSON decode error: Expecting ',' delimiter: line 7 column 287 (char 393)
Problematic caption: end_header_id|>

{ 
"Title": "Cute Owl T-Shirt", 
...
JSON data not found in caption: end_header_id|>

I cannot provide a response as it...
JSON data not found in caption: end_header_id|>

**Item Description**

*   **Title...
JSON data not found in caption: end_header_id|>

I cannot help with that request.<...
JSON data not found in caption: end_header_id|>

I'm unable to assist with that re...
JSON data not found in caption: end_header_id|>

**Product Description**

* **Titl...
JSON data not found in caption: end_header_id|>

**Product Description**

* Title:...
JSON data not found in caption: end_header_id|>

{"Title": "Ladies' Formal Jacket"...
JSON data not found in caption: end_header_id|>

Here is a rephrased version of th...
JSON data not found in caption: end_header_id|>

Here is the caption in the format...
JSON data not found in caption: end_header_id|>

**Dictionary Format Caption**

* ...
JSON data not found in caption: end_header_id|>

**Product Description**

{"Title"...
JSON data not found in caption: end_header_id|>

I can't help but feel like I've g...
JSON data not found in caption: end_header_id|>

{
  "Title": "Women's Grey Pants"...
JSON decode error: Expecting ',' delimiter: line 7 column 162 (char 272)
Problematic caption: end_header_id|>

{ 
"Title": "Anna Montanara Slipp...
JSON data not found in caption: end_header_id|>

Here is the description of the cl...
JSON data not found in caption: end_header_id|>

{ "Title": "Cycling Shorts", "Siz...
JSON decode error: Expecting ',' delimiter: line 1 column 406 (char 405)
Problematic caption: end_header_id|>

{ "Title": "Formal Pants with Zip...
JSON data not found in caption: end_header_id|>

I can't confidently answer this q...
JSON data not found in caption: end_header_id|>

**Description of a White T-Shirt ...
JSON decode error: Expecting ',' delimiter: line 1 column 408 (char 407)
Problematic caption: end_header_id|>

{"Title": "Grey Sequin Cat T-Shir...
JSON data not found in caption: end_header_id|>

Here is the caption for the image...
JSON data not found in caption: end_header_id|>

Here is the description of the cl...
JSON data not found in caption: end_header_id|>

Here is a caption for the image i...
JSON decode error: Expecting ',' delimiter: line 7 column 114 (char 226)
Problematic caption: end_header_id|>

{ 
"Title": "Mountain Hiking T-Sh...
---------------------------------------------------------------------------
KeyError                                  Traceback (most recent call last)
File ~/.conda/envs/final-checking-meta/lib/python3.12/site-packages/pandas/core/indexes/base.py:3805, in Index.get_loc(self, key)
   3804 try:
-> 3805     return self._engine.get_loc(casted_key)
   3806 except KeyError as err:

File index.pyx:167, in pandas._libs.index.IndexEngine.get_loc()

File index.pyx:196, in pandas._libs.index.IndexEngine.get_loc()

File pandas/_libs/hashtable_class_helper.pxi:7081, in pandas._libs.hashtable.PyObjectHashTable.get_item()

File pandas/_libs/hashtable_class_helper.pxi:7089, in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 'Filename'

The above exception was the direct cause of the following exception:

KeyError                                  Traceback (most recent call last)
Cell In[33], line 27
     25     # Fill NaN values with empty strings
     26     metadata = metadata.apply(lambda x: {k: v if v is not None else '' for k, v in x.items()})
---> 27     df = pd.concat([df['Filename'], pd.DataFrame(metadata.tolist())], axis=1)
     28     dataframes.append(df)
     30 # Concatenate all dataframes

File ~/.conda/envs/final-checking-meta/lib/python3.12/site-packages/pandas/core/frame.py:4102, in DataFrame.__getitem__(self, key)
   4100 if self.columns.nlevels > 1:
   4101     return self._getitem_multilevel(key)
-> 4102 indexer = self.columns.get_loc(key)
   4103 if is_integer(indexer):
   4104     indexer = [indexer]

File ~/.conda/envs/final-checking-meta/lib/python3.12/site-packages/pandas/core/indexes/base.py:3812, in Index.get_loc(self, key)
   3807     if isinstance(casted_key, slice) or (
   3808         isinstance(casted_key, abc.Iterable)
   3809         and any(isinstance(x, slice) for x in casted_key)
   3810     ):
   3811         raise InvalidIndexError(key)
-> 3812     raise KeyError(key) from err
   3813 except TypeError:
   3814     # If we have a listlike key, _check_indexing_error will raise
   3815     #  InvalidIndexError. Otherwise we fall through and re-raise
   3816     #  the TypeError.
   3817     self._check_indexing_error(key)

KeyError: 'Filename'

Verifica la diferencia de limpieza:

len(result) - result['Title'].isna().sum()
np.int64(3117)
result['Title'].describe()
count                 3117
unique                2757
top       Blue Denim Jeans
freq                    16
Name: Title, dtype: object
result
Filename  \
0     d7ed1d64-2c65-427f-9ae4-eb4aaa3e2389.jpg   
1     5c1b7a77-1fa3-4af8-9722-cd38e45d89da.jpg   
2     b2e084c7-e3a0-4182-8671-b908544a7cf2.jpg   
3     9d053b67-64e1-4050-a509-27332b9eca54.jpg   
4     d885f493-1070-4d51-bd11-f1ec156a2aa7.jpg   
...                                        ...   
5751  ae9cec7a-dd1d-49bc-adae-6446429c03d8.jpg   
5752  de853711-0b97-45a6-a794-3c424246db03.jpg   
5753  d4b0b957-5632-4df1-aba6-e562e2a84687.jpg   
5754  89074ff2-ebfe-4790-892e-8513625a05b0.jpg   
5755  0949e8e0-c807-4b6d-8453-80a05f1b733e.jpg   

                                                  Title Size Category  Gender  \
0     Stylish and Trendy Tank Top with Celestial Design    M     Tops       F   
1                              Classic White Sweatshirt    M     Tops       F   
2                                          Grey T-shirt    M  T-Shirt  Unisex   
3                                                   NaN  NaN      NaN     NaN   
4                                                   NaN  NaN      NaN     NaN   
...                                                 ...  ...      ...     ...   
5751  Men's Light Blue and White Striped Long-Sleeve...    M     Tops       M   
5752                                     Black Sneakers    S    Shoes       U   
5753                 Gray T-Shirt with Hood and Graphic    M  T-Shirt       M   
5754                                                NaN  NaN      NaN     NaN   
5755                                                NaN  NaN      NaN     NaN   

        Type                                        Description size  
0     Casual  This white tank top is a stylish and trendy pi...  NaN  
1     Casual  This classic white sweatshirt is a timeless pi...  NaN  
2     Casual  This is a short-sleeved, crew neck t-shirt tha...  NaN  
3        NaN                                                NaN  NaN  
4        NaN                                                NaN  NaN  
...      ...                                                ...  ...  
5751  Casual  This men's light blue and white striped long-s...  NaN  
5752  Casual  These sleek and versatile black sneakers are a...  NaN  
5753  Casual  The gray t-shirt with a hood and graphic is a ...  NaN  
5754     NaN                                                NaN  NaN  
5755     NaN                                                NaN  NaN  

[5756 rows x 8 columns]
Filename Title Size Category Gender Type Description size
0 d7ed1d64-2c65-427f-9ae4-eb4aaa3e2389.jpg Stylish and Trendy Tank Top with Celestial Design M Tops F Casual This white tank top is a stylish and trendy pi... NaN
1 5c1b7a77-1fa3-4af8-9722-cd38e45d89da.jpg Classic White Sweatshirt M Tops F Casual This classic white sweatshirt is a timeless pi... NaN
2 b2e084c7-e3a0-4182-8671-b908544a7cf2.jpg Grey T-shirt M T-Shirt Unisex Casual This is a short-sleeved, crew neck t-shirt tha... NaN
3 9d053b67-64e1-4050-a509-27332b9eca54.jpg NaN NaN NaN NaN NaN NaN NaN
4 d885f493-1070-4d51-bd11-f1ec156a2aa7.jpg NaN NaN NaN NaN NaN NaN NaN
... ... ... ... ... ... ... ... ... ...
5751 ae9cec7a-dd1d-49bc-adae-6446429c03d8.jpg Men's Light Blue and White Striped Long-Sleeve... M Tops M Casual This men's light blue and white striped long-s... NaN
5752 de853711-0b97-45a6-a794-3c424246db03.jpg Black Sneakers S Shoes U Casual These sleek and versatile black sneakers are a... NaN
5753 d4b0b957-5632-4df1-aba6-e562e2a84687.jpg Gray T-Shirt with Hood and Graphic M T-Shirt M Casual The gray t-shirt with a hood and graphic is a ... NaN
5754 89074ff2-ebfe-4790-892e-8513625a05b0.jpg NaN NaN NaN NaN NaN NaN NaN
5755 0949e8e0-c807-4b6d-8453-80a05f1b733e.jpg NaN NaN NaN NaN NaN NaN NaN

5756 rows × 8 columns

Eliminemos los ejemplos NaN y la columna size. Fuimos bastante ambiciosos al añadir un filtro de tamaño cuando empezamos a construir el ejemplo RAG. Ahora, esta es otra tarea para el lector que eliminamos:

# Remove rows with NaN in the 'Description' column
result = result.dropna(subset=['Description'])

# Remove the final column ('size')
result = result.drop(columns=['size'])
---------------------------------------------------------------------------
KeyError                                  Traceback (most recent call last)
Cell In[43], line 5
      2 result = result.dropna(subset=['Description'])
      4 # Remove the final column ('size')
----> 5 result = result.drop(columns=['size'])
      7 # Display the first few rows of the cleaned DataFrame
      8 result.head()

File ~/.conda/envs/final-checking-meta/lib/python3.12/site-packages/pandas/core/frame.py:5581, in DataFrame.drop(self, labels, axis, index, columns, level, inplace, errors)
   5433 def drop(
   5434     self,
   5435     labels: IndexLabel | None = None,
   (...)
   5442     errors: IgnoreRaise = "raise",
   5443 ) -> DataFrame | None:
   5444     """
   5445     Drop specified labels from rows or columns.
   5446 
   (...)
   5579             weight  1.0     0.8
   5580     """
-> 5581     return super().drop(
   5582         labels=labels,
   5583         axis=axis,
   5584         index=index,
   5585         columns=columns,
   5586         level=level,
   5587         inplace=inplace,
   5588         errors=errors,
   5589     )

File ~/.conda/envs/final-checking-meta/lib/python3.12/site-packages/pandas/core/generic.py:4788, in NDFrame.drop(self, labels, axis, index, columns, level, inplace, errors)
   4786 for axis, labels in axes.items():
   4787     if labels is not None:
-> 4788         obj = obj._drop_axis(labels, axis, level=level, errors=errors)
   4790 if inplace:
   4791     self._update_inplace(obj)

File ~/.conda/envs/final-checking-meta/lib/python3.12/site-packages/pandas/core/generic.py:4830, in NDFrame._drop_axis(self, labels, axis, level, errors, only_slice)
   4828         new_axis = axis.drop(labels, level=level, errors=errors)
   4829     else:
-> 4830         new_axis = axis.drop(labels, errors=errors)
   4831     indexer = axis.get_indexer(new_axis)
   4833 # Case for non-unique axis
   4834 else:

File ~/.conda/envs/final-checking-meta/lib/python3.12/site-packages/pandas/core/indexes/base.py:7070, in Index.drop(self, labels, errors)
   7068 if mask.any():
   7069     if errors != "ignore":
-> 7070         raise KeyError(f"{labels[mask].tolist()} not found in axis")
   7071     indexer = indexer[~mask]
   7072 return self.delete(indexer)

KeyError: "['size'] not found in axis"
result.head()
Filename  \
0  d7ed1d64-2c65-427f-9ae4-eb4aaa3e2389.jpg   
1  5c1b7a77-1fa3-4af8-9722-cd38e45d89da.jpg   
2  b2e084c7-e3a0-4182-8671-b908544a7cf2.jpg   
5  87846aa9-86cc-404a-af2c-7e8fe941081d.jpg   
7  04fa06fb-d71a-4293-9804-fe799375a682.jpg   

                                               Title Size  Category  Gender  \
0  Stylish and Trendy Tank Top with Celestial Design    M      Tops       F   
1                           Classic White Sweatshirt    M      Tops       F   
2                                       Grey T-shirt    M   T-Shirt  Unisex   
5                          Long-Sleeved V-Neck Shirt    L      Tops       U   
7                     Silver Metallic Buckle Sandals    L  Footwear       F   

     Type                                        Description  
0  Casual  This white tank top is a stylish and trendy pi...  
1  Casual  This classic white sweatshirt is a timeless pi...  
2  Casual  This is a short-sleeved, crew neck t-shirt tha...  
5  Casual  A long-sleeved, V-neck shirt with a solid purp...  
7  Casual  These silver metallic buckle sandals feature a...
Filename Title Size Category Gender Type Description
0 d7ed1d64-2c65-427f-9ae4-eb4aaa3e2389.jpg Stylish and Trendy Tank Top with Celestial Design M Tops F Casual This white tank top is a stylish and trendy pi...
1 5c1b7a77-1fa3-4af8-9722-cd38e45d89da.jpg Classic White Sweatshirt M Tops F Casual This classic white sweatshirt is a timeless pi...
2 b2e084c7-e3a0-4182-8671-b908544a7cf2.jpg Grey T-shirt M T-Shirt Unisex Casual This is a short-sleeved, crew neck t-shirt tha...
5 87846aa9-86cc-404a-af2c-7e8fe941081d.jpg Long-Sleeved V-Neck Shirt L Tops U Casual A long-sleeved, V-neck shirt with a solid purp...
7 04fa06fb-d71a-4293-9804-fe799375a682.jpg Silver Metallic Buckle Sandals L Footwear F Casual These silver metallic buckle sandals feature a...
print("\nCategory Counts:")
print(result['Category'].value_counts())
Category Counts:
Category
Tops                   1259
T-Shirt                 514
Pants                   386
Shoes                   173
Jeans                   160
Shorts                  129
Skirts                  118
Footwear                 79
Dress                    73
Jacket                   39
Coat                     21
Shirts                   17
Jackets                  17
Dresses                  16
Top                      11
Hats                      9
Skirt                     9
T-Shirts                  8
Headwear                  7
Shirt                     6
Coats                     6
Vest                      6
Jumpsuit                  5
Sweaters                  5
Accessories               4
Caps                      3
Hat                       3
Headgear                  3
Onesies                   3
Hats and Caps             3
Casual Wear               2
Denim                     2
Bottoms                   2
Bodysuit                  1
Pants and Tops            1
Sleepwear                 1
Legwear                   1
Swimwear                  1
Pants and Jackets         1
Bodysuits                 1
Jackets and Blazers       1
Casual                    1
Jumpsuits                 1
Work Pants                1
Pouf                      1
Bathrobe                  1
Tights                    1
Blazers                   1
Swimsuits                 1
Sweater                   1
T-shirt                   1
Sweatshirts               1
Name: count, dtype: int64
print("\nType Counts:")
print(result['Type'].value_counts())
Type Counts:
Type
Casual         2754
Formal          208
Lounge          128
Work Casual      15
Workout           3
Footwear          2
Athletic          2
Swimming          1
Work              1
Sleepwear         1
Home Decor        1
Swimwear          1
Name: count, dtype: int64

El modelo todavía alucina y se desvía con algunas categorías; arreglemos esto reasignándolas:

def map_category(category):
    category = category.lower()
    if 'shirt' in category or 'top' in category:
        return 'T-Shirt' if 't-shirt' in category else 'Tops'
    elif 'shoe' in category or 'footwear' in category:
        return 'Shoes'
    elif 'pant' in category:
        return 'Pants'
    elif 'jean' in category:
        return 'Jeans'
    elif 'short' in category:
        return 'Shorts'
    elif 'skirt' in category:
        return 'Skirts'
    else:
        return 'Other'

# Apply the mapping function to the 'Category' column
result['New_Category'] = result['Category'].apply(map_category)

# Print the distribution of new categories
print("Distribution of New Categories:")
print(result['New_Category'].value_counts())

# Print the mapping of old categories to new categories
print("\nMapping of Old Categories to New Categories:")
print(result.groupby('Category')['New_Category'].first().sort_index())
Distribution of New Categories:
New_Category
Tops       1295
T-Shirt     523
Pants       388
Shoes       252
Other       243
Jeans       160
Shorts      129
Skirts      127
Name: count, dtype: int64

Mapping of Old Categories to New Categories:
Category
Accessories              Other
Bathrobe                 Other
Blazers                  Other
Bodysuit                 Other
Bodysuits                Other
Bottoms                  Other
Caps                     Other
Casual                   Other
Casual Wear              Other
Coat                     Other
Coats                    Other
Denim                    Other
Dress                    Other
Dresses                  Other
Footwear                 Shoes
Hat                      Other
Hats                     Other
Hats and Caps            Other
Headgear                 Other
Headwear                 Other
Jacket                   Other
Jackets                  Other
Jackets and Blazers      Other
Jeans                    Jeans
Jumpsuit                 Other
Jumpsuits                Other
Legwear                  Other
Onesies                  Other
Pants                    Pants
Pants and Jackets        Pants
Pants and Tops            Tops
Pouf                     Other
Shirt                     Tops
Shirts                    Tops
Shoes                    Shoes
Shorts                  Shorts
Skirt                   Skirts
Skirts                  Skirts
Sleepwear                Other
Sweater                  Other
Sweaters                 Other
Sweatshirts               Tops
Swimsuits                Other
Swimwear                 Other
T-Shirt                T-Shirt
T-Shirts               T-Shirt
T-shirt                T-Shirt
Tights                   Other
Top                       Tops
Tops                      Tops
Vest                     Other
Work Pants               Pants
Name: New_Category, dtype: object

También podemos reasignar las categorías de la siguiente manera:

def map_type(type_):
    type_ = type_.lower()
    if type_ in ['casual', 'workout', 'athletic', 'swimming', 'swimwear', 'footwear']:
        return 'Casual'
    elif type_ in ['formal', 'work casual', 'work']:
        return 'Formal'
    elif type_ in ['lounge', 'sleepwear', 'home decor']:
        return 'Lounge'
    else:
        return 'Casual'  # Default to Casual for any unmatched types

# Apply the mapping function to the 'Type' column
result['New_Type'] = result['Type'].apply(map_type)

# Print the distribution of new types
print("Distribution of New Types:")
print(result['New_Type'].value_counts())

# Print the mapping of old types to new types
print("\nMapping of Old Types to New Types:")
print(result.groupby('Type')['New_Type'].first().sort_index())
Distribution of New Types:
New_Type
Casual    2763
Formal     224
Lounge     130
Name: count, dtype: int64

Mapping of Old Types to New Types:
Type
Athletic       Casual
Casual         Casual
Footwear       Casual
Formal         Formal
Home Decor     Lounge
Lounge         Lounge
Sleepwear      Lounge
Swimming       Casual
Swimwear       Casual
Work           Formal
Work Casual    Formal
Workout        Casual
Name: New_Type, dtype: object
plt.style.use('ggplot')

# Create a figure with two subplots
fig, (ax1, ax2) = plt.subplots(nrows=2, ncols=1, figsize=(12, 16))

# Plot distribution of Categories
sns.countplot(data=result, y='Category', ax=ax1, order=result['New_Category'].value_counts().index)
ax1.set_title('Distribution of Categories', fontsize=16)
ax1.set_xlabel('Count', fontsize=12)
ax1.set_ylabel('Category', fontsize=12)

# Plot distribution of Types
sns.countplot(data=result, y='Type', ax=ax2, order=result['New_Type'].value_counts().index)
ax2.set_title('Distribution of Types', fontsize=16)
ax2.set_xlabel('Count', fontsize=12)
ax2.set_ylabel('Type', fontsize=12)

# Adjust layout and display the plot
plt.tight_layout()
plt.show()

# Optional: Save the figure
# plt.savefig('category_type_distribution.png', dpi=300, bbox_inches='tight')

# Additional analysis: Print top 5 categories and types
print("Top 5 Categories:")
print(result['Category'].value_counts().head())

print("\nTop 5 Types:")
print(result['Type'].value_counts().head())
<Figure size 1200x1600 with 2 Axes>
Top 5 Categories:
Category
Tops       1259
T-Shirt     514
Pants       386
Shoes       173
Jeans       160
Name: count, dtype: int64

Top 5 Types:
Type
Casual         2754
Formal          208
Lounge          128
Work Casual      15
Workout           3
Name: count, dtype: int64
def sample_category(group):
    if len(group) > 100:
        return group.sample(n=100, random_state=42)
    else:
        return group

# Group by New_Category and apply the sampling function
sampled_data = result.groupby('New_Category').apply(sample_category).reset_index(drop=True)

# Print the distribution of categories in the sampled data
print("Distribution of Categories in Sampled Data:")
print(sampled_data['New_Category'].value_counts())

# Print the distribution of types in the sampled data
print("\nDistribution of Types in Sampled Data:")
print(sampled_data['New_Type'].value_counts())

# Calculate and print percentages
total = len(sampled_data)
print("\nPercentage Distribution of Categories:")
category_percentage = (sampled_data['New_Category'].value_counts() / total * 100).round(2)
print(category_percentage)

print("\nPercentage Distribution of Types:")
type_percentage = (sampled_data['New_Type'].value_counts() / total * 100).round(2)
print(type_percentage)

# Print the total number of items in the sampled dataset
print(f"\nTotal number of items in the sampled dataset: {len(sampled_data)}")
Distribution of Categories in Sampled Data:
New_Category
Jeans      100
Other      100
Pants      100
Shoes      100
Shorts     100
Skirts     100
T-Shirt    100
Tops       100
Name: count, dtype: int64

Distribution of Types in Sampled Data:
New_Type
Casual    700
Formal     64
Lounge     36
Name: count, dtype: int64

Percentage Distribution of Categories:
New_Category
Jeans      12.5
Other      12.5
Pants      12.5
Shoes      12.5
Shorts     12.5
Skirts     12.5
T-Shirt    12.5
Tops       12.5
Name: count, dtype: float64

Percentage Distribution of Types:
New_Type
Casual    87.5
Formal     8.0
Lounge     4.5
Name: count, dtype: float64

Total number of items in the sampled dataset: 800
/tmp/ipykernel_525083/1300003174.py:8: DeprecationWarning: DataFrameGroupBy.apply operated on the grouping columns. This behavior is deprecated, and in a future version of pandas the grouping columns will be excluded from the operation. Either pass `include_groups=False` to exclude the groupings or explicitly select the grouping columns after groupby to silence this warning.
  sampled_data = result.groupby('New_Category').apply(sample_category).reset_index(drop=True)

Ahora podemos volver a muestrear y tener un conjunto de datos agradable y equilibrado:

def sample_category(group):
    if len(group) > 100:
        return group.sample(n=100, random_state=42)
    else:
        return group

# Group by New_Category and apply the sampling function
sampled_data = result.groupby('New_Category').apply(sample_category).reset_index(drop=True)

# Set up the matplotlib figure
fig, axs = plt.subplots(2, 2, figsize=(20, 15))

# 1. Bar plot of Category distribution
sns.countplot(data=sampled_data, x='New_Category', order=sampled_data['New_Category'].value_counts().index, ax=axs[0, 0])
axs[0, 0].set_title('Distribution of Categories')
axs[0, 0].set_xticklabels(axs[0, 0].get_xticklabels(), rotation=45, ha='right')

# 2. Bar plot of Type distribution
sns.countplot(data=sampled_data, x='New_Type', order=sampled_data['New_Type'].value_counts().index, ax=axs[0, 1])
axs[0, 1].set_title('Distribution of Types')
axs[0, 1].set_xticklabels(axs[0, 1].get_xticklabels(), rotation=45, ha='right')

# 3. Heatmap of Category vs Type
cross_tab = pd.crosstab(sampled_data['New_Category'], sampled_data['New_Type'])
sns.heatmap(cross_tab, annot=True, fmt='d', cmap='YlGnBu', ax=axs[1, 0])
axs[1, 0].set_title('Heatmap of Category vs Type')

# 4. Grouped bar plot of Type distribution within each Category
cross_tab_normalized = cross_tab.div(cross_tab.sum(axis=1), axis=0)
cross_tab_normalized.plot(kind='bar', stacked=False, ax=axs[1, 1])
axs[1, 1].set_title('Type Distribution within each Category')
axs[1, 1].set_xlabel('Category')
axs[1, 1].set_ylabel('Proportion')
axs[1, 1].legend(title='Type', bbox_to_anchor=(1.05, 1), loc='upper left')
axs[1, 1].set_xticklabels(axs[1, 1].get_xticklabels(), rotation=45, ha='right')

# Adjust layout and display the plot
plt.tight_layout()
plt.show()

# Print the total number of items in the sampled dataset
print(f"Total number of items in the sampled dataset: {len(sampled_data)}")
/tmp/ipykernel_525083/3643476101.py:8: DeprecationWarning: DataFrameGroupBy.apply operated on the grouping columns. This behavior is deprecated, and in a future version of pandas the grouping columns will be excluded from the operation. Either pass `include_groups=False` to exclude the groupings or explicitly select the grouping columns after groupby to silence this warning.
  sampled_data = result.groupby('New_Category').apply(sample_category).reset_index(drop=True)
/tmp/ipykernel_525083/3643476101.py:16: UserWarning: set_ticklabels() should only be used with a fixed number of ticks, i.e. after set_ticks() or using a FixedLocator.
  axs[0, 0].set_xticklabels(axs[0, 0].get_xticklabels(), rotation=45, ha='right')
/tmp/ipykernel_525083/3643476101.py:21: UserWarning: set_ticklabels() should only be used with a fixed number of ticks, i.e. after set_ticks() or using a FixedLocator.
  axs[0, 1].set_xticklabels(axs[0, 1].get_xticklabels(), rotation=45, ha='right')
<Figure size 2000x1500 with 5 Axes>
Total number of items in the sampled dataset: 800
result.head()
Filename  \
0  d7ed1d64-2c65-427f-9ae4-eb4aaa3e2389.jpg   
1  5c1b7a77-1fa3-4af8-9722-cd38e45d89da.jpg   
2  b2e084c7-e3a0-4182-8671-b908544a7cf2.jpg   
5  87846aa9-86cc-404a-af2c-7e8fe941081d.jpg   
7  04fa06fb-d71a-4293-9804-fe799375a682.jpg   

                                               Title Size  Category  Gender  \
0  Stylish and Trendy Tank Top with Celestial Design    M      Tops       F   
1                           Classic White Sweatshirt    M      Tops       F   
2                                       Grey T-shirt    M   T-Shirt  Unisex   
5                          Long-Sleeved V-Neck Shirt    L      Tops       U   
7                     Silver Metallic Buckle Sandals    L  Footwear       F   

     Type                                        Description New_Category  \
0  Casual  This white tank top is a stylish and trendy pi...         Tops   
1  Casual  This classic white sweatshirt is a timeless pi...         Tops   
2  Casual  This is a short-sleeved, crew neck t-shirt tha...      T-Shirt   
5  Casual  A long-sleeved, V-neck shirt with a solid purp...         Tops   
7  Casual  These silver metallic buckle sandals feature a...        Shoes   

  New_Type  
0   Casual  
1   Casual  
2   Casual  
5   Casual  
7   Casual
Filename Title Size Category Gender Type Description New_Category New_Type
0 d7ed1d64-2c65-427f-9ae4-eb4aaa3e2389.jpg Stylish and Trendy Tank Top with Celestial Design M Tops F Casual This white tank top is a stylish and trendy pi... Tops Casual
1 5c1b7a77-1fa3-4af8-9722-cd38e45d89da.jpg Classic White Sweatshirt M Tops F Casual This classic white sweatshirt is a timeless pi... Tops Casual
2 b2e084c7-e3a0-4182-8671-b908544a7cf2.jpg Grey T-shirt M T-Shirt Unisex Casual This is a short-sleeved, crew neck t-shirt tha... T-Shirt Casual
5 87846aa9-86cc-404a-af2c-7e8fe941081d.jpg Long-Sleeved V-Neck Shirt L Tops U Casual A long-sleeved, V-neck shirt with a solid purp... Tops Casual
7 04fa06fb-d71a-4293-9804-fe799375a682.jpg Silver Metallic Buckle Sandals L Footwear F Casual These silver metallic buckle sandals feature a... Shoes Casual
final_data = result.drop(columns=['Type', 'Category'])

# Rename 'New_Type' to 'Type' and 'New_Category' to 'Category'
final_data = final_data.rename(columns={'New_Type': 'Type', 'New_Category': 'Category'})

# Print the first few rows of the final dataset
print("\nFirst few rows of the final dataset:")
print(final_data.head())

# Print the column names of the final dataset
print("\nColumns in the final dataset:")
print(final_data.columns.tolist())

# Save the final DataFrame
final_data.to_csv('final_balanced_sample_dataset.csv', index=False)
First few rows of the final dataset:
                                   Filename  \
0  d7ed1d64-2c65-427f-9ae4-eb4aaa3e2389.jpg   
1  5c1b7a77-1fa3-4af8-9722-cd38e45d89da.jpg   
2  b2e084c7-e3a0-4182-8671-b908544a7cf2.jpg   
5  87846aa9-86cc-404a-af2c-7e8fe941081d.jpg   
7  04fa06fb-d71a-4293-9804-fe799375a682.jpg   

                                               Title Size  Gender  \
0  Stylish and Trendy Tank Top with Celestial Design    M       F   
1                           Classic White Sweatshirt    M       F   
2                                       Grey T-shirt    M  Unisex   
5                          Long-Sleeved V-Neck Shirt    L       U   
7                     Silver Metallic Buckle Sandals    L       F   

                                         Description Category    Type  
0  This white tank top is a stylish and trendy pi...     Tops  Casual  
1  This classic white sweatshirt is a timeless pi...     Tops  Casual  
2  This is a short-sleeved, crew neck t-shirt tha...  T-Shirt  Casual  
5  A long-sleeved, V-neck shirt with a solid purp...     Tops  Casual  
7  These silver metallic buckle sandals feature a...    Shoes  Casual  

Columns in the final dataset:
['Filename', 'Title', 'Size', 'Gender', 'Description', 'Category', 'Type']

Final dataset saved as 'final_balanced_sample_dataset.csv'

Siguiente paso

¡Hemos avanzado mucho! Ahora nuestro conjunto de datos es excelente para ser incrustado y utilizado en nuestro paso final.

La siguiente parte será la más fácil; sin embargo, todavía haremos un poco de ingeniería de prompt.

#fin
Lección del curso «Llama Cookbook (use cases)» de Meta, publicado con licencia MIT. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Meta. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios
← AnteriorSiguiente: Parte 3: Configuración del ejemplo RAG y validación del pipeline de recuperación →