Etiquetar y subtitular imágenes de ropa
Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
Usarás las capacidades de visión del modelo Gemini y el modelo de embeddings para añadir etiquetas y subtítulos a imágenes de prendas de vestir.
Estas descripciones se pueden usar junto con los embeddings para permitirte buscar prendas de vestir específicas usando lenguaje natural u otras imágenes.
Configuración
%pip install -U -q "google-genai>=2.9.0"
from google import genai
from google.genai import types
Configura tu clave de API
Para ejecutar la siguiente celda, tu clave de API debe estar almacenada en un Secreto de Colab llamado GEMINI_API_KEY. Si aún no tienes una clave de API, o no estás seguro de cómo crear un Secreto de Colab, consulta Autenticación para ver un ejemplo.
from google.colab import userdata
api_key = userdata.get('GEMINI_API_KEY')
client = genai.Client(api_key=api_key)
Descargando el conjunto de datos
Primero, necesitas descargar un conjunto de datos con imágenes. Contiene imágenes de varias prendas de vestir que puedes usar para probar el modelo.
!wget https://storage.googleapis.com/generativeai-downloads/data/clothes-dataset.zip
--2025-04-08 18:09:36-- https://storage.googleapis.com/generativeai-downloads/data/clothes-dataset.zip
Resolving storage.googleapis.com (storage.googleapis.com)... 172.217.218.207, 142.251.31.207, 142.251.18.207, ...
Connecting to storage.googleapis.com (storage.googleapis.com)|172.217.218.207|:443... connected.
HTTP request sent, awaiting response... 200 OK
Length: 730831 (714K) [application/zip]
Saving to: ‘clothes-dataset.zip.3’
clothes-dataset.zip 100%[===================>] 713.70K 1.42MB/s in 0.5s
2025-04-08 18:09:37 (1.42 MB/s) - ‘clothes-dataset.zip.3’ saved [730831/730831]
Descomprime los datos en clothes-dataset.zip y colócalos en una carpeta en tu entorno de Colab.
!unzip -o clothes-dataset.zip
Archive: clothes-dataset.zip
inflating: clothes-dataset/6.jpg
inflating: clothes-dataset/4.jpg
inflating: clothes-dataset/1.jpg
inflating: clothes-dataset/2.jpg
inflating: clothes-dataset/7.jpg
inflating: clothes-dataset/9.jpg
inflating: clothes-dataset/8.jpg
inflating: clothes-dataset/10.jpg
inflating: clothes-dataset/5.jpg
inflating: clothes-dataset/3.jpg
from glob import glob
images = glob("/content/clothes-dataset/*")
images.sort(reverse=True)
Generando palabras clave
Puedes usar el LLM para extraer palabras clave relevantes de las imágenes.
Aquí tienes una función de ayuda para llamar a la API de Gemini con imágenes. El "sleep" es para asegurar que no se exceda la cuota. Consulta nuestra página de precios para conocer las cuotas actuales.
from PIL import Image as PILImage
import time
MODEL_ID = 'gemini-3.7-flash' # @param ['gemini-3.1-pro-preview', 'gemini-3.7-flash', 'gemini-3.5-flash-lite', 'gemini-2.5-pro'] {"allow-input":true, isTemplate: true}
# a helper function for calling
def generate_text_using_image(prompt, image_path, sleep_time=4):
start = time.perf_counter()
response = client.models.generate_content(
model=MODEL_ID,
contents=[PILImage.open(image_path)],
config=types.GenerateContentConfig(
system_instruction=prompt
),
)
end = time.perf_counter()
duration = end - start
time.sleep(sleep_time - duration if duration < sleep_time else 0)
return response.text
Primero, define la lista de posibles palabras clave.
import numpy as np
keywords = np.concatenate((
["flannel", "shorts", "pants", "dress", "T-shirt", "shirt", "suit"],
["women", "men", "boys", "girls"],
["casual", "sport", "elegant"],
["fall", "winter", "spring", "summer"],
["red", "violet", "blue", "green", "yellow", "orange", "black", "white"],
["polyester", "cotton", "denim", "silk", "leather", "wool", "fur"]
)
)
Ahora, define un prompt que te ayudará a definir palabras clave que describan la ropa. En el siguiente prompt, se utiliza el "few-shot prompting" para preparar al LLM con ejemplos de cómo se deben generar estas palabras clave y cuáles son válidas.
keyword_prompt = f"""
You are an expert in clothing that specializes in tagging images of clothes,
shoes, and accessories.
Your job is to extract all relevant keywords from
a photo that will help describe an item.
You are going to see an image,
extract only the keywords for the clothing, and try to provide as many
keywords as possible.
Allowed keywords: {list(keywords)}
Extract tags only when it is obvious that it describes the main item in
the image. Return the keywords as a list of strings:
example1: ["blue", "shoes", "denim"]
example2: ["sport", "skirt", "cotton", "blue", "red"]
"""
def generate_keywords(image_path):
return generate_text_using_image(keyword_prompt, image_path)
Genera palabras clave para cada una de las imágenes.
from IPython.display import Image, display
for image_path in images[:4]:
response_text = generate_keywords(image_path)
display(Image(image_path))
print(response_text)
<IPython.core.display.Image object>
["shorts", "denim", "blue"]
<IPython.core.display.Image object>
["suit", "men", "blue", "elegant"]
<IPython.core.display.Image object>
["suit", "blue", "black", "men", "elegant"]
<IPython.core.display.Image object>
Here are the extracted keywords:
["T-shirt", "cotton", "casual", "women", "spring", "summer", "red"]
Corrección y deduplicación de palabras clave
Desafortunadamente, a pesar de proporcionar una lista de posibles palabras clave, el modelo, al menos en teoría, puede devolver una palabra clave no válida. Puede ser un duplicado, por ejemplo, "denim" para "jeans", o ser completamente ajena a cualquier palabra clave de la lista.
Para abordar estos problemas, puedes usar embeddings para mapear las palabras clave a las predefinidas y eliminar las no relacionadas.
import pandas as pd
EMBEDDINGS_MODEL_ID = "embedding-001" # @param ["gemini-embedding-exp-03-07", "gemini-embedding-2-preview", "embedding-001"] {"allow-input":true, isTemplate: true}
def embed(text):
embedding = client.models.embed_content(
model=EMBEDDINGS_MODEL_ID,
contents=text,
config=types.EmbedContentConfig(
task_type="semantic_similarity"
)
)
return np.array(embedding.embeddings[0].values)
keywords_df = pd.DataFrame({'Keywords': keywords})
keywords_df["Embeddings"] = keywords_df['Keywords'].apply(embed)
keywords_df.head()
Keywords Embeddings
0 flannel [-0.010662257, -0.036850516, 0.041096948, -0.0...
1 shorts [0.033239827, 0.00013335898, 0.031834923, -0.0...
2 pants [0.012273628, -0.024099143, 0.0032160978, -0.0...
3 dress [0.010586792, -0.023003917, -0.028866835, -0.0...
4 T-shirt [-0.015854606, -0.055863477, 0.025282638, -0.0...
| Keywords | Embeddings | |
|---|---|---|
| 0 | flannel | [-0.010662257, -0.036850516, 0.041096948, -0.0... |
| 1 | shorts | [0.033239827, 0.00013335898, 0.031834923, -0.0... |
| 2 | pants | [0.012273628, -0.024099143, 0.0032160978, -0.0... |
| 3 | dress | [0.010586792, -0.023003917, -0.028866835, -0.0... |
| 4 | T-shirt | [-0.015854606, -0.055863477, 0.025282638, -0.0... |
Para fines de demostración, define una función que evalúe la similitud entre dos vectores de embedding. En este caso, usarás la similitud del coseno, pero otras medidas como el producto escalar también funcionan.
def cosine_similarity(array_1, array_2):
return np.dot(array_1,array_2)/(np.linalg.norm(array_1)*np.linalg.norm(array_2))
A continuación, define una función que te permita reemplazar una palabra clave con la palabra más similar en el dataframe de palabras clave que creaste previamente.
Ten en cuenta que el umbral se decide arbitrariamente, puede requerir ajustes dependiendo del caso de uso y del conjunto de datos.
def replace_word_with_most_similar(keyword, keywords_df, threshold=0.7):
# No need for embeddings if the keyword is valid.
if keyword in keywords_df["Keywords"]:
return keyword
embedding = embed(keyword)
similarities = keywords_df['Embeddings'].apply(lambda row_embedding: cosine_similarity(embedding, row_embedding))
most_similar_keyword_index = similarities.idxmax()
if similarities[most_similar_keyword_index] < threshold:
return None
return keywords_df.loc[most_similar_keyword_index, "Keywords"]
Aquí tienes un ejemplo de cómo estas palabras clave se pueden mapear a una palabra clave con el significado más cercano.
for word in ["purple", "tank top", "everyday"]:
print(word, "->", replace_word_with_most_similar(word, keywords_df))
purple -> violet
tank top -> T-shirt
everyday -> casual
Ahora puedes dejar las palabras que no encajan en nuestras categorías predefinidas o eliminarlas. En este escenario, todas las palabras sin un reemplazo adecuado se omitirán.
def map_generated_keywords_to_predefined(generated_keywords, keywords_df=keywords_df):
output_keywords = set()
for keyword in generated_keywords:
if mapped_keyword := replace_word_with_most_similar(keyword, keywords_df):
output_keywords.add(mapped_keyword)
return output_keywords
print(map_generated_keywords_to_predefined(["white", "business", "sport", "women", "polyester"]))
print(map_generated_keywords_to_predefined(["blue", "jeans", "women", "denim", "casual"]))
{'polyester', 'women', 'white', 'sport'}
{'women', 'blue', 'casual', 'denim'}
Generando subtítulos
caption_prompt ="""
You are an expert in clothing that specializes in describing images of
clothes, shoes and accessories.
Your job is to extract information from a photo that will help describe an item.
You are going to see an image, focus only on the piece of clothing,
ignore suroundings.
Be specific, but stay concise, the description should only be one sentence long.
Most important aspects are color, type of clothing, material, style
and who is it meant for.
If you are not sure about a part of the image, ignore it.
"""
def generate_caption(image_path):
return generate_text_using_image(caption_prompt, image_path)
for image_path in images[8:]:
response_text = generate_caption(image_path)
display(Image(image_path))
print(response_text)
<IPython.core.display.Image object>
This is a red, short-sleeved, knee-length women's dress with a colorful floral pattern.
<IPython.core.display.Image object>
This is a khaki button-up shirt with two chest pockets and long sleeves, designed for men.
Buscando ropa específica
Preparando nuestro conjunto de datos
Primero, necesitas generar un subtítulo y palabras clave para cada imagen. Luego, usarás embeddings, que se usarán más tarde para comparar las imágenes en el conjunto de datos de búsqueda con otras descripciones e imágenes.
Además, la función de ayuda ast.literal_eval() te permite evaluar un objeto pasado y obtener el objeto literal. Por ejemplo, si pasaras una cadena "[1, 2, 3]", la función ast.literal_eval() la devolvería como una lista [1, 2, 3]. Para obtener más información sobre esta función, aquí tienes la documentación.
import ast
def generate_keyword_and_caption(image_path):
keywords = generate_keywords(image_path)
try:
keywords = ast.literal_eval(keywords)
keywords = map_generated_keywords_to_predefined(keywords)
except SyntaxError:
pass
caption = generate_caption(image_path)
return {
"image_path": image_path,
"keywords": keywords,
"caption": caption
}
Usarás solo las primeras 8 imágenes, así que el resto se puede usar para probar.
described_df = pd.DataFrame([generate_keyword_and_caption(image_path) for image_path in images[:8]])
def embed_row(row):
text = ", ".join(row["keywords"]) + ".\n" + row["caption"]
return embed(text)
described_df["embeddings"] = described_df.apply(lambda x: embed_row(x), axis=1)
described_df
image_path \
0 /content/clothes-dataset/9.jpg
1 /content/clothes-dataset/8.jpg
2 /content/clothes-dataset/7.jpg
3 /content/clothes-dataset/6.jpg
4 /content/clothes-dataset/5.jpg
5 /content/clothes-dataset/4.jpg
6 /content/clothes-dataset/3.jpg
7 /content/clothes-dataset/2.jpg
keywords \
0 {blue, shorts, denim}
1 {men, elegant, blue, suit}
2 {black, blue, suit}
3 {T-shirt, casual, shirt, red, women, cotton}
4 {dress, violet}
5 {dress, women, summer}
6 {blue, pants, denim}
7 {flannel, shirt}
caption \
0 These are blue denim shorts with distressed de...
1 This is a men's formal blue suit with a collar...
2 This is a men's blue suit jacket with a black ...
3 This is an oversized women's t-shirt dress in ...
4 This is a wine-colored dress with a v-neck, sh...
5 This is a colorful, short-sleeved dress with a...
6 This is a light blue pair of denim jeans for a...
7 This is a long-sleeved, button-up flannel shir...
embeddings
0 [0.059082124, 0.0006507264, 0.039219383, -0.06...
1 [0.03221182, -0.045238137, -0.030657658, -0.03...
2 [0.04739029, -0.08439743, -0.0054425937, -0.03...
3 [0.034199536, -0.04078525, -0.0031016248, -0.0...
4 [0.03572897, -0.030525232, 0.0014781117, -0.04...
5 [0.06510386, -0.011200937, -0.0020332306, -0.0...
6 [0.055783305, -0.03525318, 0.0073569883, -0.06...
7 [-0.0145449005, -0.01969576, 0.045668166, -0.0...
| image_path | keywords | caption | embeddings | |
|---|---|---|---|---|
| 0 | /content/clothes-dataset/9.jpg | {blue, shorts, denim} | These are blue denim shorts with distressed de... | [0.059082124, 0.0006507264, 0.039219383, -0.06... |
| 1 | /content/clothes-dataset/8.jpg | {men, elegant, blue, suit} | This is a men's formal blue suit with a collar... | [0.03221182, -0.045238137, -0.030657658, -0.03... |
| 2 | /content/clothes-dataset/7.jpg | {black, blue, suit} | This is a men's blue suit jacket with a black ... | [0.04739029, -0.08439743, -0.0054425937, -0.03... |
| 3 | /content/clothes-dataset/6.jpg | {T-shirt, casual, shirt, red, women, cotton} | This is an oversized women's t-shirt dress in ... | [0.034199536, -0.04078525, -0.0031016248, -0.0... |
| 4 | /content/clothes-dataset/5.jpg | {dress, violet} | This is a wine-colored dress with a v-neck, sh... | [0.03572897, -0.030525232, 0.0014781117, -0.04... |
| 5 | /content/clothes-dataset/4.jpg | {dress, women, summer} | This is a colorful, short-sleeved dress with a... | [0.06510386, -0.011200937, -0.0020332306, -0.0... |
| 6 | /content/clothes-dataset/3.jpg | {blue, pants, denim} | This is a light blue pair of denim jeans for a... | [0.055783305, -0.03525318, 0.0073569883, -0.06... |
| 7 | /content/clothes-dataset/2.jpg | {flannel, shirt} | This is a long-sleeved, button-up flannel shir... | [-0.0145449005, -0.01969576, 0.045668166, -0.0... |