useful tips for datasets

Just go to the link and download the dataset and look at how it was prepared.

https://civitai.com/models/2771332/krea2-john-william-waterhouse

AiToolKit no longer requires square images but you want all your image with clean aspect ratios.

1:1, 1:2, 1:3, 2:3, 1:4

And you want each edge to be a multiple of 256.

---

https://www.youtube.com/watch?v=OCsqHdHf81M

In this video he shows you how to use ChatGPT to do the caption. I do not recommend using his markdown file though.

---

Copy and paste and save this...it's my captioning system. Just paste it into ChatGPT and then upload zip files of the images.

Just change the first two sentences to match your dataset.

Gil Elvgren LoRA Captioning Instructions

You are captioning a ZIP file containing pin-up artwork by Gil Elvgren for LoRA training.

Required Header

Begin every caption with this exact line:

Pin-up painting, style of Gil Elvgren, gilelvg

Add one blank line after the header.

Core Accuracy Rule

Caption only details that are visibly present in the individual image.

Do not infer missing clothing, shoes, jewelry, stockings, garters, props, materials, colors, patterns, hairstyles, makeup, nail polish, or background objects.

When a detail is ambiguous, omit it.

Never complete an outfit based on what would normally match the theme.

Examples:

Do not add heels when the feet are outside the frame.
Do not call fabric silk, satin, leather, lace, or chiffon unless the material is visually identifiable.
Do not add stockings merely because garters are present.
Do not add garters merely because stockings are present.
Do not assume red nail polish unless the nails are visible and clearly red.
If gloves cover the hands or fingers, omit the Nails line unless the nails are still clearly visible.
Do not assume earrings, bracelets, necklaces, hats, gloves, or hair accessories.
Do not infer an object from the general scene when it cannot be clearly identified.

It is better to omit one uncertain detail than to add one incorrect detail.

Individual-Image Workflow

Do not caption from contact sheets.

Open and inspect every original image separately at full available resolution.

Complete the batch using this process:

  1. Open one individual image.

  2. Inspect the entire image.

  3. Inspect the face, hair, hands, clothing, legs, feet, props, and background separately.

  4. Write the caption for that image.

  5. Inspect the same image a second time.

  6. Verify every line of the caption against the image.

  7. Amend or remove any unsupported line.

  8. Save the caption as a matching .txt file.

  9. Continue to the next individual image.

  10. Zip all completed .txt captions only after the entire batch has been verified.

Create a working folder for the captions and keep the completed files there until the batch is finished.

Verification Standard

During the second inspection, check every caption line with these questions:

Is every noun in this line visibly present?
Is the color accurate?
Is the garment type accurate?
Is the material clearly identifiable?
Is the pattern actually visible?
Is the body orientation correct?
Are the correct arms and legs described?
Is the object held by the correct hand?
Are the feet visible?
Are shoes actually visible?
Are stockings or garter bands clearly visible?
Is the hairstyle described from visible structure rather than assumption?
Are the nails directly visible?
Are gloves obscuring the nails?
Is each background object identifiable?
Did I add a conventional pin-up detail that is not actually shown?

Delete or simplify any line that fails verification.

Required Sections

Use these sections when applicable:

Concept
Pose
Attire
Hair Makeup Nails
Expression
Background

Props may be included as a separate section when the image contains several important handheld or scene objects.

Do not include an empty section.

Concept Section

Write one concrete sentence describing the basic visible scene.

Prioritize the woman, her action, the main prop, and the setting.

Avoid subjective or conceptual language such as:

glamorous
seductive
luxurious
enchanting
playful atmosphere
elegant mood
cinematic
dramatic beauty

Use concrete descriptions instead.

Example:

Blonde woman seated on a wooden ladder while holding several books in a library

Pose Section

Write one concrete pose detail per line.

Use separate lines rather than a paragraph.

Example:

Pose
Full body front-facing view
Standing with legs apart
Left hand resting on hip
Right hand holding a paintbrush
Torso angled slightly right
Head tilted slightly left
Eyes looking toward viewer

Describe only what is visible.

Useful pose details include:

full body
three-quarter body
front view
side view
three-quarter back view
back view
seated
standing
kneeling
reclining
bending forward
leaning backward
weight resting on one leg
legs crossed at knees
legs crossed at ankles
one knee raised
one arm extended
hand resting on hip
head turned over shoulder
gaze direction

Do not confuse overlapping legs with crossed legs.

Attire Formatting

Use one line for each visible garment or accessory category.

Combine all details about the same item on one line using commas.

Correct:

Dress: white summer dress, fitted bodice, plunge neckline, short puff sleeves, full knee-length skirt, scalloped lace hem
Panties: pale pink high-waisted panties, glossy sheen, dark blue floral embroidery at hips
Stockings: sheer black nylon thigh-high stockings, wide opaque garter bands, visible back seams
Heels: black closed-toe pumps, pointed toes, slender high heels
Gloves: white wrist-length gloves

Incorrect:

Dress: white dress
Dress: fitted bodice
Dress: short sleeves
Dress: lace hem

Do not repeat identical labels on multiple lines.

Use specific category names when visible:

Dress
Blouse
Shirt
Top
Bra
Corset
Bodice
Jacket
Skirt
Shorts
Pants
Panties
Garter belt
Garters
Stockings
Socks
Shoes
Heels
Boots
Gloves
Hat
Scarf
Belt
Necklace
Earrings
Bracelet
Robe
Apron
Swimsuit
Bikini top
Bikini bottoms

Nudity and Bare Feet

When no clothing is visible on the upper body, use:

Upper body: nude

When no clothing is visible on the lower body, use:

Lower body: nude

When both are visible and nude, include both lines.

When the feet are visible and no shoes or socks are worn, use:

Feet: bare

When the feet are outside the frame or obscured, omit footwear entirely.

Do not write barefoot unless the bare feet are actually visible.

Hair Makeup Nails Formatting

Use one consolidated line for each category.

Correct:

Hair: blonde shoulder-length hair, large curled waves, fringe bangs
Makeup: dark eyeliner, long lashes, blue eyeshadow, pink blush, glossy red lipstick
Nails: red nail polish

Incorrect:

Hair: blonde hair
Hair: shoulder-length hair
Hair: curled waves

Incorrect:

Makeup: dark eyeliner
Makeup: long lashes
Makeup: pink blush
Makeup: red lipstick

Describe only visible features.

Possible hair details:

hair color
approximate length
straight
wavy
curled
ringlets
victory rolls
rolled bangs
fringe bangs
side part
center part
ponytail
bun
updo
loose curls
hair ribbon
flower accessory

Do not identify a hairstyle as victory rolls unless the rolled structure is clearly visible.

Possible makeup details:

dark eyeliner
winged eyeliner
long lashes
blue eyeshadow
green eyeshadow
pink blush
red lipstick
pink lipstick
glossy lipstick
defined brows

Omit makeup details that cannot be resolved from the image.

For nails use:

Nails: red nail polish

Do not use:

Nails: red manicure

Omit the Nails line when the nails are not clearly visible.

If gloves cover the hands or fingers, omit the Nails line unless the nails remain directly visible.

Never caption nail polish through gloves.

Expression Section

Use concrete facial observations.

Examples:

Expression
Wide-eyed surprised expression
Raised eyebrows
Rounded open mouth forming an oh shape
Eyes looking toward viewer

Expression
Broad smile
Eyes looking toward viewer

Expression
Focused expression
Eyes looking downward toward the book

Avoid interpreting emotions beyond visible facial features.

Background Section

Caption only identifiable visible background elements.

Use one object or closely related group per line.

Example:

Background
Tall wooden bookshelves filled with books
Wooden library ladder
Several books falling through the air
Pale wooden floor

Do not describe lighting unless it is an important visible element requested by the user.

Do not add generic environmental objects to make the scene feel complete.

Props Section

Use a Props section when several important objects are interacting with the subject.

Example:

Props
Open black suitcase filled with clothing
Small black dog pulling a garment from the suitcase
Red travel tag attached to suitcase handle

Only describe identifiable objects.

Language Style

Use concrete nouns and restrained adjectives.

Good:

red full skirt
wooden chair
black dog
white towel
round hand mirror
sheer black stockings
curled blonde hair
open suitcase
metal ladder
blue wall

Avoid unnecessary aesthetic terms:

gorgeous
sultry
alluring
luxurious
romantic
dreamy
elegant
sophisticated
captivating
glamorous

Do not mention artistic technique, brushwork, composition quality, or painterly atmosphere beyond the required header.

File Handling

Each image receives one matching .txt file.

Preserve the image filename exactly, changing only the extension to .txt.

Example:

GilElvgren (31).jpg

becomes:

GilElvgren (31).txt

If two images share the same stem but have different extensions, add a short extension identifier so neither caption is overwritten.

Place all completed caption files in one folder.

After every image has been individually inspected and every caption has been verified against its original image a second time, zip the .txt files and provide the ZIP for download.

Final Quality Rule

Accuracy is more important than caption length.

A shorter caption containing only verified details is better than a detailed caption containing one hallucinated item.

 

https://www.reddit.com/r/StableDiffusion/comments/1v4we4h/the_wonders_of_krea2/

This article was updated on July 26, 2026