Lección 33 · 10 min · Gratis

Etiquetar y subtitular imágenes de ropa

Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

Usarás las capacidades de visión del modelo Gemini y el modelo de embeddings para añadir etiquetas y subtítulos a imágenes de prendas de vestir.

Estas descripciones se pueden usar junto con los embeddings para permitirte buscar prendas de vestir específicas usando lenguaje natural u otras imágenes.

Configuración

%pip install -U -q "google-genai>=2.9.0"
from google import genai
from google.genai import types

Configura tu clave de API

Para ejecutar la siguiente celda, tu clave de API debe estar almacenada en un Secreto de Colab llamado GEMINI_API_KEY. Si aún no tienes una clave de API, o no estás seguro de cómo crear un Secreto de Colab, consulta Autenticación para ver un ejemplo.

from google.colab import userdata
api_key = userdata.get('GEMINI_API_KEY')

client = genai.Client(api_key=api_key)

Descargando el conjunto de datos

Primero, necesitas descargar un conjunto de datos con imágenes. Contiene imágenes de varias prendas de vestir que puedes usar para probar el modelo.

!wget https://storage.googleapis.com/generativeai-downloads/data/clothes-dataset.zip
--2025-04-08 18:09:36--  https://storage.googleapis.com/generativeai-downloads/data/clothes-dataset.zip
Resolving storage.googleapis.com (storage.googleapis.com)... 172.217.218.207, 142.251.31.207, 142.251.18.207, ...
Connecting to storage.googleapis.com (storage.googleapis.com)|172.217.218.207|:443... connected.
HTTP request sent, awaiting response... 200 OK
Length: 730831 (714K) [application/zip]
Saving to: ‘clothes-dataset.zip.3’

clothes-dataset.zip 100%[===================>] 713.70K  1.42MB/s    in 0.5s    

2025-04-08 18:09:37 (1.42 MB/s) - ‘clothes-dataset.zip.3’ saved [730831/730831]

Descomprime los datos en clothes-dataset.zip y colócalos en una carpeta en tu entorno de Colab.

!unzip -o clothes-dataset.zip
Archive:  clothes-dataset.zip
  inflating: clothes-dataset/6.jpg   
  inflating: clothes-dataset/4.jpg   
  inflating: clothes-dataset/1.jpg   
  inflating: clothes-dataset/2.jpg   
  inflating: clothes-dataset/7.jpg   
  inflating: clothes-dataset/9.jpg   
  inflating: clothes-dataset/8.jpg   
  inflating: clothes-dataset/10.jpg  
  inflating: clothes-dataset/5.jpg   
  inflating: clothes-dataset/3.jpg
from glob import glob
images = glob("/content/clothes-dataset/*")
images.sort(reverse=True)

Generando palabras clave

Puedes usar el LLM para extraer palabras clave relevantes de las imágenes.

Aquí tienes una función de ayuda para llamar a la API de Gemini con imágenes. El "sleep" es para asegurar que no se exceda la cuota. Consulta nuestra página de precios para conocer las cuotas actuales.

from PIL import Image as PILImage
import time

MODEL_ID = 'gemini-3.7-flash' # @param ['gemini-3.1-pro-preview', 'gemini-3.7-flash', 'gemini-3.5-flash-lite', 'gemini-2.5-pro'] {"allow-input":true, isTemplate: true}
# a helper function for calling

def generate_text_using_image(prompt, image_path, sleep_time=4):
  start = time.perf_counter()
  response = client.models.generate_content(
    model=MODEL_ID,
    contents=[PILImage.open(image_path)],
    config=types.GenerateContentConfig(
        system_instruction=prompt
    ),
)
  end = time.perf_counter()
  duration = end - start
  time.sleep(sleep_time - duration if duration < sleep_time else 0)
  return response.text

Primero, define la lista de posibles palabras clave.

import numpy as np
keywords = np.concatenate((
    ["flannel", "shorts", "pants", "dress", "T-shirt", "shirt", "suit"],
    ["women", "men", "boys", "girls"],
    ["casual", "sport", "elegant"],
    ["fall", "winter", "spring", "summer"],
    ["red", "violet", "blue", "green", "yellow", "orange", "black", "white"],
    ["polyester", "cotton", "denim", "silk", "leather", "wool", "fur"]
)
)

Ahora, define un prompt que te ayudará a definir palabras clave que describan la ropa. En el siguiente prompt, se utiliza el "few-shot prompting" para preparar al LLM con ejemplos de cómo se deben generar estas palabras clave y cuáles son válidas.

keyword_prompt = f"""
     You are an expert in clothing that specializes in tagging images of clothes,
     shoes, and accessories.
     Your job is to extract all relevant keywords from
     a photo that will help describe an item.
     You are going to see an image,
     extract only the keywords for the clothing, and try to provide as many
     keywords as possible.

     Allowed keywords: {list(keywords)}

     Extract tags only when it is obvious that it describes the main item in
     the image. Return the keywords as a list of strings:

     example1: ["blue", "shoes", "denim"]
     example2: ["sport", "skirt", "cotton", "blue", "red"]
"""
def generate_keywords(image_path):
  return generate_text_using_image(keyword_prompt, image_path)

Genera palabras clave para cada una de las imágenes.

from IPython.display import Image, display
for image_path in images[:4]:
  response_text = generate_keywords(image_path)
  display(Image(image_path))
  print(response_text)
<IPython.core.display.Image object>
["shorts", "denim", "blue"]
<IPython.core.display.Image object>
["suit", "men", "blue", "elegant"]
<IPython.core.display.Image object>
["suit", "blue", "black", "men", "elegant"]
<IPython.core.display.Image object>
Here are the extracted keywords:
["T-shirt", "cotton", "casual", "women", "spring", "summer", "red"]

Corrección y deduplicación de palabras clave

Desafortunadamente, a pesar de proporcionar una lista de posibles palabras clave, el modelo, al menos en teoría, puede devolver una palabra clave no válida. Puede ser un duplicado, por ejemplo, "denim" para "jeans", o ser completamente ajena a cualquier palabra clave de la lista.

Para abordar estos problemas, puedes usar embeddings para mapear las palabras clave a las predefinidas y eliminar las no relacionadas.

import pandas as pd

EMBEDDINGS_MODEL_ID = "embedding-001" # @param ["gemini-embedding-exp-03-07", "gemini-embedding-2-preview", "embedding-001"] {"allow-input":true, isTemplate: true}

def embed(text):
    embedding = client.models.embed_content(
        model=EMBEDDINGS_MODEL_ID,
        contents=text,
        config=types.EmbedContentConfig(
            task_type="semantic_similarity"
        )
    )
    return np.array(embedding.embeddings[0].values)

keywords_df = pd.DataFrame({'Keywords': keywords})
keywords_df["Embeddings"] = keywords_df['Keywords'].apply(embed)
keywords_df.head()
Keywords                                         Embeddings
0  flannel  [-0.010662257, -0.036850516, 0.041096948, -0.0...
1   shorts  [0.033239827, 0.00013335898, 0.031834923, -0.0...
2    pants  [0.012273628, -0.024099143, 0.0032160978, -0.0...
3    dress  [0.010586792, -0.023003917, -0.028866835, -0.0...
4  T-shirt  [-0.015854606, -0.055863477, 0.025282638, -0.0...
Keywords Embeddings
0 flannel [-0.010662257, -0.036850516, 0.041096948, -0.0...
1 shorts [0.033239827, 0.00013335898, 0.031834923, -0.0...
2 pants [0.012273628, -0.024099143, 0.0032160978, -0.0...
3 dress [0.010586792, -0.023003917, -0.028866835, -0.0...
4 T-shirt [-0.015854606, -0.055863477, 0.025282638, -0.0...

Para fines de demostración, define una función que evalúe la similitud entre dos vectores de embedding. En este caso, usarás la similitud del coseno, pero otras medidas como el producto escalar también funcionan.

def cosine_similarity(array_1, array_2):
  return np.dot(array_1,array_2)/(np.linalg.norm(array_1)*np.linalg.norm(array_2))

A continuación, define una función que te permita reemplazar una palabra clave con la palabra más similar en el dataframe de palabras clave que creaste previamente.

Ten en cuenta que el umbral se decide arbitrariamente, puede requerir ajustes dependiendo del caso de uso y del conjunto de datos.

def replace_word_with_most_similar(keyword, keywords_df, threshold=0.7):
  # No need for embeddings if the keyword is valid.
  if keyword in keywords_df["Keywords"]:
    return keyword
  embedding = embed(keyword)
  similarities = keywords_df['Embeddings'].apply(lambda row_embedding: cosine_similarity(embedding, row_embedding))
  most_similar_keyword_index = similarities.idxmax()
  if similarities[most_similar_keyword_index] < threshold:
    return None
  return keywords_df.loc[most_similar_keyword_index, "Keywords"]

Aquí tienes un ejemplo de cómo estas palabras clave se pueden mapear a una palabra clave con el significado más cercano.

for word in ["purple", "tank top", "everyday"]:
  print(word, "->", replace_word_with_most_similar(word, keywords_df))
purple -> violet
tank top -> T-shirt
everyday -> casual

Ahora puedes dejar las palabras que no encajan en nuestras categorías predefinidas o eliminarlas. En este escenario, todas las palabras sin un reemplazo adecuado se omitirán.

def map_generated_keywords_to_predefined(generated_keywords, keywords_df=keywords_df):
  output_keywords = set()
  for keyword in generated_keywords:
    if mapped_keyword := replace_word_with_most_similar(keyword, keywords_df):
      output_keywords.add(mapped_keyword)
  return output_keywords

print(map_generated_keywords_to_predefined(["white", "business", "sport", "women", "polyester"]))
print(map_generated_keywords_to_predefined(["blue", "jeans", "women", "denim", "casual"]))
{'polyester', 'women', 'white', 'sport'}
{'women', 'blue', 'casual', 'denim'}

Generando subtítulos

caption_prompt ="""
     You are an expert in clothing that specializes in describing images of
     clothes, shoes and accessories.
     Your job is to extract information from a photo that will help describe an item.
     You are going to see an image, focus only on the piece of clothing,
     ignore suroundings.
     Be specific, but stay concise, the description should only be one sentence long.
     Most important aspects are color, type of clothing, material, style
     and who is it meant for.
     If you are not sure about a part of the image, ignore it.
"""
def generate_caption(image_path):
  return generate_text_using_image(caption_prompt, image_path)

for image_path in images[8:]:
  response_text = generate_caption(image_path)
  display(Image(image_path))
  print(response_text)
<IPython.core.display.Image object>
This is a red, short-sleeved, knee-length women's dress with a colorful floral pattern.
<IPython.core.display.Image object>
This is a khaki button-up shirt with two chest pockets and long sleeves, designed for men.

Buscando ropa específica

Preparando nuestro conjunto de datos

Primero, necesitas generar un subtítulo y palabras clave para cada imagen. Luego, usarás embeddings, que se usarán más tarde para comparar las imágenes en el conjunto de datos de búsqueda con otras descripciones e imágenes.

Además, la función de ayuda ast.literal_eval() te permite evaluar un objeto pasado y obtener el objeto literal. Por ejemplo, si pasaras una cadena "[1, 2, 3]", la función ast.literal_eval() la devolvería como una lista [1, 2, 3]. Para obtener más información sobre esta función, aquí tienes la documentación.

import ast

def generate_keyword_and_caption(image_path):
  keywords = generate_keywords(image_path)
  try:
    keywords = ast.literal_eval(keywords)
    keywords = map_generated_keywords_to_predefined(keywords)
  except SyntaxError:
    pass
  caption = generate_caption(image_path)
  return {
      "image_path": image_path,
      "keywords": keywords,
      "caption": caption
  }

Usarás solo las primeras 8 imágenes, así que el resto se puede usar para probar.

described_df = pd.DataFrame([generate_keyword_and_caption(image_path) for image_path in images[:8]])
def embed_row(row):
  text = ", ".join(row["keywords"]) + ".\n" + row["caption"]
  return embed(text)
described_df["embeddings"] = described_df.apply(lambda x: embed_row(x), axis=1)
described_df
image_path  \
0  /content/clothes-dataset/9.jpg   
1  /content/clothes-dataset/8.jpg   
2  /content/clothes-dataset/7.jpg   
3  /content/clothes-dataset/6.jpg   
4  /content/clothes-dataset/5.jpg   
5  /content/clothes-dataset/4.jpg   
6  /content/clothes-dataset/3.jpg   
7  /content/clothes-dataset/2.jpg   

                                       keywords  \
0                         {blue, shorts, denim}   
1                    {men, elegant, blue, suit}   
2                           {black, blue, suit}   
3  {T-shirt, casual, shirt, red, women, cotton}   
4                               {dress, violet}   
5                        {dress, women, summer}   
6                          {blue, pants, denim}   
7                              {flannel, shirt}   

                                             caption  \
0  These are blue denim shorts with distressed de...   
1  This is a men's formal blue suit with a collar...   
2  This is a men's blue suit jacket with a black ...   
3  This is an oversized women's t-shirt dress in ...   
4  This is a wine-colored dress with a v-neck, sh...   
5  This is a colorful, short-sleeved dress with a...   
6  This is a light blue pair of denim jeans for a...   
7  This is a long-sleeved, button-up flannel shir...   

                                          embeddings  
0  [0.059082124, 0.0006507264, 0.039219383, -0.06...  
1  [0.03221182, -0.045238137, -0.030657658, -0.03...  
2  [0.04739029, -0.08439743, -0.0054425937, -0.03...  
3  [0.034199536, -0.04078525, -0.0031016248, -0.0...  
4  [0.03572897, -0.030525232, 0.0014781117, -0.04...  
5  [0.06510386, -0.011200937, -0.0020332306, -0.0...  
6  [0.055783305, -0.03525318, 0.0073569883, -0.06...  
7  [-0.0145449005, -0.01969576, 0.045668166, -0.0...
image_path keywords caption embeddings
0 /content/clothes-dataset/9.jpg {blue, shorts, denim} These are blue denim shorts with distressed de... [0.059082124, 0.0006507264, 0.039219383, -0.06...
1 /content/clothes-dataset/8.jpg {men, elegant, blue, suit} This is a men's formal blue suit with a collar... [0.03221182, -0.045238137, -0.030657658, -0.03...
2 /content/clothes-dataset/7.jpg {black, blue, suit} This is a men's blue suit jacket with a black ... [0.04739029, -0.08439743, -0.0054425937, -0.03...
3 /content/clothes-dataset/6.jpg {T-shirt, casual, shirt, red, women, cotton} This is an oversized women's t-shirt dress in ... [0.034199536, -0.04078525, -0.0031016248, -0.0...
4 /content/clothes-dataset/5.jpg {dress, violet} This is a wine-colored dress with a v-neck, sh... [0.03572897, -0.030525232, 0.0014781117, -0.04...
5 /content/clothes-dataset/4.jpg {dress, women, summer} This is a colorful, short-sleeved dress with a... [0.06510386, -0.011200937, -0.0020332306, -0.0...
6 /content/clothes-dataset/3.jpg {blue, pants, denim} This is a light blue pair of denim jeans for a... [0.055783305, -0.03525318, 0.0073569883, -0.06...
7 /content/clothes-dataset/2.jpg {flannel, shirt} This is a long-sleeved, button-up flannel shir... [-0.0145449005, -0.01969576, 0.045668166, -0.0...

Encontrando ropa usando lenguaje natural

def find_image_from_text(text):
  text_embedding = embed(text)
  similarities = described_df['embeddings'].apply(lambda row_embedding: cosine_similarity(text_embedding, row_embedding))
  most_fitting_image_index = similarities.idxmax()
  return described_df["image_path"][most_fitting_image_index]
display(Image(find_image_from_text("A suit for a wedding.")))
<IPython.core.display.Image object>
display(Image(find_image_from_text("A colorful dress.")))
<IPython.core.display.Image object>

Encontrando ropa similar usando imágenes

def find_image_from_image(image_path):
  text_embedding = embed_row(generate_keyword_and_caption(image_path))
  similarities = described_df['embeddings'].apply(lambda row_embedding: cosine_similarity(text_embedding, row_embedding))
  most_fitting_image_index = similarities.idxmax()
  return described_df["image_path"][most_fitting_image_index]
image_path = images[8]
display(Image(image_path))
display(Image(find_image_from_image(image_path)))
<IPython.core.display.Image object>
<IPython.core.display.Image object>
image_path = images[9]
display(Image(image_path))
display(Image(find_image_from_image(image_path)))
<IPython.core.display.Image object>
<IPython.core.display.Image object>

Resumen

Has usado el SDK de Python de la API de Gemini para etiquetar y subtitular imágenes de ropa. Usando modelos de embedding, pudiste buscar en una base de datos de imágenes ropa que coincidiera con nuestra descripción, o similar a la prenda proporcionada.

Lección del curso «Gemini API Cookbook (examples)» de Google, publicado con licencia Apache 2.0. Traducción y adaptación al español de IA con Clase. IA con Clase no está afiliado a Google. Ver el original · Licencia
Esta lección es gratuita. El resto del curso se abre con la Membresía de IA con Clase, que incluye todos los cursos del catálogo. Ver precios
← AnteriorSiguiente: Habla con documentos usando embeddings →