Hugging Face Authors Build Six Models With ML-intern
- •Hugging Face authors describe building six custom models with ML-intern over several days, with total compute charges of about USD 103.
- •Citrus model accuracy rose from 14.9% to 52.8% across 335 test photos, at about USD 1.90 compute cost.
- •Agate student reached 0.536 GenEval in four steps, against 0.563 for the teacher at 50 steps.
Hugging Face authors Yuvraj Sharma and Abubakar Abid published a guide on October 8, 2026, describing how they used ML-intern in HuggingChat to build six custom models over several days. The agent plans work, requests a budget before paid jobs, runs a small test, then trains, evaluates and publishes models using Hugging Face hardware. Across the projects listed, reported compute charges totaled about USD 103.
Sharma says prompts specify the goal, dataset, base model and training script, with verified information separated so the agent does not need to rediscover it. He recommends asking for a baseline score before training and a smoke test (a small run that checks the setup) before paying for a full job. Prompts also define deliverables and spending limits; ML-intern starts each task with a zero-dollar budget and asks permission before paid work. His first citrus-model prompt was about 450 words; by his sixth project, prompts were closer to 2,000 words.
For a citrus disease model, the team combined three Project-AgML datasets into 3,017 annotated images covering 21 pests, illnesses, nutritional gaps and treatments. Fine-tuning Qwen3.5-2B raised the correct-problem rate on 335 test photos from the base model’s 14.9% to 52.8% after two epochs on one A10G, at about USD 1.90. A separate Huggy LoRA (a small add-on trained to modify a model) used 84 captioned drawings with FLUX.2 klein base 4B; step 200 was the first checkpoint judged fully on-model, while from step 500 the style began appearing in unrelated prompts. Its compute cost was about USD 7.60.
The Viewpoint Orbit camera-angle LoRA rendered 1,030 household objects at 24 angles, creating 24,722 transparent images. Training used 461 objects and 1,844 before-and-after pairs across 23 camera instructions; 40 objects were held out for testing. It ran 2,000 steps in about 90 minutes on one A100, and the project cost about USD 16.
The Doodle-in LoRA dataset paired photos with removed objects and magenta replacement marks. It contained 6,042 training pairs and 160 test pairs, including 40 examples from 23 object classes excluded from training. After 2,000 steps in 1 hour 38 minutes on one A100, the selected step-500 checkpoint, paired with Viggle turbo LoRA at 6 steps, placed objects where marked in 67.5% of cases, taking 4.7 seconds per edit. Results for unseen classes were 65.0%, compared with 64.2% for the other classes. The project cost about USD 24.
Pocket Rewriter distilled a 9B prompt rewriter into 0.8B and 2B students using 1,840 selected examples from 8,797 generated requests. The 0.8B version returns valid output 99.7% of the time, uses about a quarter of the teacher’s tokens, and is available as an 812 MB GGUF file for CPUs. The project cost about USD 16.05.
For Agate Preview 002, a 260M-parameter text-to-image model, ML-intern reduced generation from 50 steps to 4. In the first run, it cached 155,000 training images as latents and the 4-step student beat the teacher at 4 steps on GenEval and FID; export to ONNX made it usable in a browser. A second run added 24,000 teacher-generated image pairs. GenEval improved from 0.509 to 0.536, versus 0.563 for the teacher at 50 steps, with the student using one-fourth the compute. The two runs cost about USD 37. The article’s project cost table totals about USD 103 across all six models.