The Train tab needs to let users caption small datasets in the browser and
pull in a ready-made set to see training work end to end, neither of which
the upload-only endpoint supported.
Add, under /api/train/diffusion/dataset:
- GET {name}/images lists every image with its resolved caption (metadata
beats a per-image sidecar, matching the trainer's discovery order) so
uncaptioned images are visible and flaggable.
- GET {name}/image/{filename} serves an image, with ?thumb=<px> returning a
cached downscaled JPEG kept in a hidden .thumbs subdir (regenerated when
the source is newer) so the labeling grid stays light.
- PUT {name}/caption/{filename} writes, or when blank clears, the .txt
sidecar; DELETE {name}/image/{filename} removes the image plus its
sidecars and thumbnails.
- GET dataset-examples lists a curated, license-labelled registry, and
POST dataset/import-example materializes one into a dataset folder as
numbered images + .txt captions. Two loaders cover the shapes seen in the
wild: streaming rows from datasets.load_dataset (dog-example, Tuxemon) and
a snapshot + jsonl walk for imagefolder repos whose captions live in a
non-standard *.jsonl (the public-domain tarot set). Imports are idempotent
and cap the image count.
Filenames and dataset names are validated against path traversal and pinned
inside the datasets root.