Discussion about this post

User's avatar
Robert Feng's avatar

Very relevant topic to what we are doing in the intersection between ai and finance.

AutomationLabs's avatar

The distinction between flexible preferences and invalidating failures is the part most eval setups quietly collapse — we ran everything on a weighted average until one hard-failure category (fabricated citations) kept passing because strong style scores averaged it out. Splitting the rubric into gating checks versus scored preferences fixed more than any reweighting did. Curious where you'd draw the line on letting a model help author its own rubric: does that widen coverage of new failure modes, or just teach it the boundaries of the test faster?

No posts

Ready for more?