How uniopen customized Amazon Nova to their retail moderation policies for production deployment
Amazon Web Services Felix Chin, Jia-You Lin
uniopen tuned Amazon Nova 2 Lite to spot retail moderation issues in its apps. The fix was a big jump in accuracy, then a small prompt tweak pushed both scores past target.
Based on reporting by Amazon Web Services, Felix Chin, Jia-You Lin — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
uniopen, the digital communication and membership platform from Taiwan’s Uni-President Enterprises Group, needed its moderation system to understand something a general model would not know on its own: the company’s own rules. The policy doesn’t just ask what happened. It also asks what the behavior refers to, with the subject bucketed as brand, other, or forbidden. That two-part call has to be right if the moderation decision is going to be useful.
So the team built a production path around Amazon Nova 2 Lite, then trained it on uniopen’s own examples in Amazon SageMaker AI. The setup kept the work split cleanly: the live moderation path stayed separate from correction, training, evaluation, and deployment. Reported mistakes could be checked by people, written into Amazon S3 as verified corrections, and then fed back into the next round of improvement. Amazon DynamoDB tracked model configurations, while Argo Workflows on Amazon EKS handled prompt optimization, evaluation, and deployment. Argo CD pushed approved changes into production.
The numbers tell the story. On a held-out test set of 737 conversation windows, the baseline Nova 2 Lite model scored 0.5852 on Per Behavior Macro F1 and 0.4162 on Subject Type Macro F1. After supervised fine-tuning with LoRA, those scores jumped to 0.8364 and 0.8302. That was already enough to show the model had learned a lot from domain-specific examples, especially on the subject-type side.
But the team didn’t stop there. A prompt-only change simplified the output from JSON to a line-based format and clarified how multiple behaviors should be returned. No retraining. That lifted the scores again, to 0.8550 for Per Behavior Macro F1 and 0.8491 for Subject Type Macro F1, clearing the production targets of 0.8500 and 0.8200.
The release process was guarded by hard gates and soft gates. Hard gates were must-pass regression tests. Soft gates flagged warning signs like low confidence or a drop in a specific class, which could hold a candidate for human review. That meant uniopen could send routine moderation down the automated path while keeping people in the loop for ambiguous content. A dull policy flow, maybe. Exactly the sort of dull that keeps production from becoming a mess.
My take — AI-written commentary, not fact-checked reporting
This is the sensible way to use foundation models: teach them the company’s rules, then make them earn their way into production. The industry loves pretending a general model is a substitute for a policy team, which is adorable until the complaints start. Human review plus hard gates is not glamorous, but it beats shipping confidence with no clue what the model is actually classifying.
Read more about this at: Amazon Web Services