Machine LearningAI Research2 min reading time

From Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers

Apple Research Blog
Read full post
Researchers developed a rubric-based reward system for open-domain question answering that uses query-specific, evidence-grounded rubrics decomposed into multiple quality dimensions. This approach improves answer quality across composition, grounding, and instruction-following by up to 6.5%. Conditioning rubrics on retrieved evidence enhances factual accuracy, while multi-dimensional rubrics improve coherence and adherence to query requirements.

More in Machine Learning

Machine Learning3 min read

Nvidia and Palantir fine-tune a 30B Nemotron model for Nvidia’s supply chain. It beats a model 18 times its size.

Covered by 3 sources
Machine Learning6 min read

CoreWeave Puts Field Engineers Inside Customer Teams for Physical AI

Covered by 2 sources
Machine Learning2 min read

Weatherwatch: AI model beats standard methods at predicting cyclones

The Guardian