seha incheon's honest review, [Chuseok Special] An LLM Evaluation Harness Engineering Guide from a Silicon Valley Developer course
Reviews 3
Average rating 5
Since there are so many lines of code and files, I felt it was a little difficult to follow along while you were proceeding by simply copying things. Still, the content was so informative and covered a topic I had never heard of before that I think I was able to tolerate those shortcomings. In particular, since the model used for evaluation cannot be trusted either, that model also needs to be evaluated. But if another model evaluates it again, then ultimately it becomes impossible to trust it. Understanding it from that perspective, I think the conclusion of model evaluation is that humans have to put in the hard work. - In fact, I understand that Meta hired more people and used more human labor to give tasks to and train models, and it seems that perspective is reflected here as well.
1
Hello, thank you for leaving such a great review!! Since there are so many evaluation methods, I was worried that covering each and every one might feel like too high a barrier to entry, so I focused primarily on what they feel like and what direction they should take. So rather than focusing on the overall code structure, I think it would be great if you learned these patterns and approaches!! I’ll continue to provide more useful content in the future!!
1

