What Do You Do When The AI Model You Need Doesn't Exist?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What Do You Do When The AI Model You Need Doesn't Exist? on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A Hugging Face contributor reports using its ML-intern agent to build and publish seven custom models over several days. The examples include a small prompt rewriter and a citrus-disease image model, but the results and costs are self-reported and have not been independently verified.

A Hugging Face contributor says the platform’s ML-intern agent helped build and publish seven custom models over several days, as described in the original account, including a small prompt rewriter and a model for identifying citrus problems in images. The account offers a practical example of agent-assisted model development, but its performance figures and compute costs are self-reported, not an independent evaluation.

The first project addressed a specific gap: the contributor wanted a smaller version of the prompt rewriter included with Qwen-Image 2.1. They said the official rewriter has 9 billion parameters, needs about 20 GB of memory and can generate thousands of tokens before producing a paragraph. The contributor found compressed versions but said they could not find a smaller alternative, so they trained a 0.8-billion-parameter model using the larger model as a teacher.

According to the contributor, the compact rewriter returned valid output 99.7% of the time and used about one-quarter as many tokens as its teacher. They put total compute costs, including having the larger model label 8,797 example requests, at about $16. The account does not provide the measurement protocol behind the validity rate, so the figure should be understood as the author’s reported result.

Another project fine-tuned Qwen3.5-2B to identify citrus pests, diseases and nutritional deficiencies from photos, a use case related to AI models for prediction. The contributor said the dataset had 3,017 annotated images covering 21 categories. On 335 test photos, they reported that accuracy on the correct problem rose from 14.9% for the base model to 52.8% after two training epochs on one A10G GPU. They reported compute costs of about $1.90.

At a glance
reportWhen: Reported in the contributor’s recent ac…
The developmentA Hugging Face contributor has published an account of using the ML-intern agent to plan, train, evaluate and publish seven custom models.
At a glance
reportWhen: Reported last week; the projects were b…
The developmentA Hugging Face contributor says an AI agent called ML-intern helped plan, train, evaluate and publish seven custom models on the Hub over several days.

A Shorter Route to Specialized Models

The account matters because it shows how an AI agent might reduce the hands-on coordination involved in adapting a model for a narrow task. Rather than personally arranging every training step, the contributor says they asked ML-intern to propose a plan, test a small run, train, evaluate and publish models using Hugging Face hardware. For people with a focused need, this could make experimentation more accessible.

The reported examples also show why lower barriers do not remove the need for careful testing. The citrus result is presented against the untuned model on the same test set, giving a comparison for that particular dataset. It does not establish how the model would perform on other orchards, cameras or conditions. Likewise, the reported compute bills cover compute costs, not a complete accounting of time spent preparing data, writing instructions or checking outputs.

Model customization can also bring unwanted behavior. The contributor said later checkpoints for a character-generation LoRA began affecting prompts that were not about the character. That account suggests training can spread a desired style too broadly, reinforcing the need to compare checkpoints and test unrelated prompts before publication.

Amazon

AI model training hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Contributor Directed ML-Intern

The contributor says each project began as a message in HuggingChat with ML-intern enabled. The agent proposed a plan and, before paid work, asked for spending approval. When a prompt did not include a budget, it reportedly offered options and asked the user to choose. The contributor says the workflow also included a small test run, training, evaluation and publication.

The instructions became more detailed over the course of the projects: the contributor says prompts grew from about 450 words on the first project to nearly 2,000 by the sixth. They included the dataset, base model and training script, along with requests to establish a baseline, run a smoke test and stay within a spending cap. One instruction asked the agent to report the base model’s score on the same metric before training, allowing the contributor to compare results.

The account also describes a character-generation LoRA trained on 84 captioned drawings and a camera-angle LoRA for Qwen-Image 2.1. For the latter, the contributor says the agent generated 24,722 transparent images of scanned household objects across 24 angles, then trained on selected objects while holding others back for testing. Training took about 90 minutes on one A100; the contributor put the whole project’s compute cost at about $16, including failed jobs that had to be resubmitted. The supplied account gives details for only some of the seven projects.

““Also report the base model’s zero-shot score on the same metric before training so we can see the gain.””

— The Hugging Face contributor

Amazon

custom AI model development kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Independent Testing Is Still Missing

The supplied account does not include independent replication or complete evaluation details for every model. The contributor’s figures are self-reported, and the material does not explain how the 99.7% valid-output rate was measured, whether image labels and test results were independently checked, or how performance changes on data outside the reported test sets.

The reported compute totals are not a full cost comparison: they do not establish the value of the contributor’s time, data preparation or review. The account also does not provide descriptions of all seven models, or comparable results from other users. As a result, it cannot show whether the reported spending and outcomes are typical of the agent’s use.

Amazon

AI image classification kit

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Published Projects Invite Follow-Up

The contributor says the models and evaluations are on Hugging Face and that all seven prompts are available in a public GitHub repository. Those materials may allow readers to inspect the project details and methods. Further checks across different users, datasets and tasks would be needed to determine how consistently the workflow produces similar results.

For future projects, the contributor’s described safeguards include recording a baseline before training, running a small smoke test, checking that saved weights changed and setting a spending limit. Whether that process is sufficient will depend on the task and the quality of the data and evaluation. No independent assessment or broader results are provided in the supplied account.

Amazon

prompt rewriting AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is ML-intern?

It is the Hugging Face agent the contributor says they used through HuggingChat to plan model projects and handle steps including test runs, training, evaluation and publication.

What models did the contributor describe?

The detailed examples include a 0.8-billion-parameter prompt rewriter, a citrus-problem image classifier, a character-generation LoRA and a camera-angle LoRA. The account says there were seven models in total but the supplied material does not detail all seven.

Are the performance and cost figures independently verified?

No independent verification is included in the supplied account. The scores and compute costs are attributed to the contributor, and some measurement details are not provided.

What did the contributor report spending?

Reported compute costs included about $1.90 for the citrus model and about $16 each for the prompt rewriter and camera-angle LoRA project. These figures are not a full accounting of data preparation, time or other possible expenses.

Primary source: Hugging Face · via ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Understanding The AI Apocalypse Through Grok And Claude’s Eyes

A Nautilus article prompts Grok and Claude to discuss AI risks, raising questions about chatbot responses and their implications for AI safety debates.

Grok 4.6: The Frontier Is Now A Price War

Grok 4.6 from SpaceXAI maintains unchanged pricing while boosting intelligence, fueling a price war at the AI frontier. Key details and implications explained.

OpenDLSS: A Vulkan Reimplementation Of Nvidia’s DLSS 5 Neural Rendering Network

OpenDLSS says its Vulkan implementation matches DLSS 5 network outputs byte for byte, but users must supply model weights and compatible hardware.