Hugging Face has published a detailed account of ML-intern, an agent mode inside HuggingChat that takes a prompt describing a model someone wants, then plans the project, asks for a spending limit before running anything, trains and evaluates the result against a baseline, and publishes it on the Hub with its evaluation in the model card.
The post, written by a Hugging Face team member, walks through six models built this way over a few days. A citrus-disease vision model, fine-tuned from Qwen3.5-2B on 3,017 annotated images, went from a 14.9% baseline accuracy to 52.8% after two epochs, for about $1.90 in compute. A 'Pocket Rewriter' — an 0.8B model distilled from a 9B prompt-rewriting model — returns valid output 99.7% of the time using a quarter of the teacher's tokens and runs on a CPU, for roughly $16. Other examples include camera-angle and doodle-replacement LoRAs for Qwen-Image 2.1, and a 4-step distilled version of a 260M-parameter text-to-image model. Total compute across all six projects came to about $103.
ML-intern begins every task at zero budget and requires explicit permission before spending, enforced through a spending cap the user sets in the prompt (for example, "cap total spend at USD 12"). The author notes that prompts improved with iteration — the first was about 450 words, the sixth closer to 2,000 — and that two elements mattered most: asking for a baseline score before training, and requiring a small smoke test before committing to a full run.
ML-intern is available now inside HuggingChat by switching on the mode and submitting a prompt; the author's own prompts are published on GitHub as starting templates.