[Chuseok Special] An LLM Evaluation Harness Engineering Guide from a Silicon Valley Developer

If you’ve ever had trouble determining whether the quality of responses declined after changing a prompt or model while building a service using an LLM, this course may be the solution. You’ll learn how to create your own evaluation criteria and automatically identify responses that have changed.

(4.8) 5 reviews

95 learners

Level Basic

Course period Unlimited

JavaScript
JavaScript
TypeScript
TypeScript
Business Productivity
Business Productivity
AI
AI
ChatGPT
ChatGPT
JavaScript
JavaScript
TypeScript
TypeScript
Business Productivity
Business Productivity
AI
AI
ChatGPT
ChatGPT

Reviews from Early Learners

4.8

5.0

Be Dev

91% enrolled

Rather than mastering an evaluation method called Eval, I think the course is more about introducing concepts and validation patterns like these. In other words, I feel it’s very helpful for broadening one’s perspective. I also thought it was quite an innovative topic while taking the course myself.

5.0

seha incheon

83% enrolled

Since there are so many lines of code and files, I felt it was a little difficult to follow along while you were proceeding by simply copying things. Still, the content was so informative and covered a topic I had never heard of before that I think I was able to tolerate those shortcomings. In particular, since the model used for evaluation cannot be trusted either, that model also needs to be evaluated. But if another model evaluates it again, then ultimately it becomes impossible to trust it. Understanding it from that perspective, I think the conclusion of model evaluation is that humans have to put in the hard work. - In fact, I understand that Meta hired more people and used more human labor to give tasks to and train models, and it seems that perspective is reflected here as well.

5.0

우당탕탕

83% enrolled

I found the lecture very informative!

What you will gain after the course

  • Implement a tool for collecting and quantitatively and qualitatively evaluating LLM responses

  • Verify quality differences before and after prompt and model changes through regression testing

  • How to Design Evaluation Datasets and Clear LLM Evaluation Criteria

  • Building a Reliable Evaluation Pipeline by Combining LLM Judges and Rule-Based Evaluation

  • Trace failed responses to analyze the cause of the problem among the prompt, retrieval, and model.

  • Connect evaluation results to CI/CD to continuously manage the quality of LLM services

Recommended for
these people

Who is this course right for?

  • Developers who operate LLM services and need to validate the impact of prompt and model changes

  • A developer who wants to systematically evaluate the response quality of LLM agents and RAG services

  • AI service developers who find it difficult to begin quality validation due to the lack of evaluation datasets and criteria

  • A developer looking to automate manual testing and introduce regression testing before deployment

  • A developer who wants to complete an LLM evaluation harness as a project that can be applied directly to real-world work.

Hello
This is Hong

Inflearn Verified

Career Verified

10,421

Learners

649

Reviews

171

Answers

4.8

Rating

34

Courses

I started studying development after becoming interested in it while idling at home, and I am currently responsible for platform server development in Pangyo. I am continuing my activities as a knowledge sharer because I want to provide you with the methods I used to study, as well as the various problems and solutions you may encounter in practice.

 

These lectures are not created solely through my own knowledge. There are others who collaborate on every lecture.

 

[Instructor Career]

[Former] Blockchain developer related to Sandbox IP

[Former] Metaverse Backend Developer

[Current] A server developer becoming a veteran in Pangyo

 

[Interview History]

[Other Inquiries]

[Official Site]

More

Curriculum

All

23 lectures ∙ (5hr 21min)

Course Materials:

Lecture resources
Published: 
Last updated: 

Reviews

All

5 reviews

4.8

5 reviews

  • lookupjoin님의 프로필 이미지
    lookupjoin

    Reviews 3

    ∙

    Average Rating 5.0

    5

    91% enrolled

    Rather than mastering an evaluation method called Eval, I think the course is more about introducing concepts and validation patterns like these. In other words, I feel it’s very helpful for broadening one’s perspective. I also thought it was quite an innovative topic while taking the course myself.

    • jhong
      Instructor

      Hello, thank you for leaving a review!! I’m glad it helped broaden your perspective!! These days, it seems that working with a broader perspective has become especially important. I’ll continue to provide more helpful content in the future!! Thank you!

  • sehaincheon1507님의 프로필 이미지
    sehaincheon1507

    Reviews 3

    ∙

    Average Rating 5.0

    5

    83% enrolled

    Since there are so many lines of code and files, I felt it was a little difficult to follow along while you were proceeding by simply copying things. Still, the content was so informative and covered a topic I had never heard of before that I think I was able to tolerate those shortcomings. In particular, since the model used for evaluation cannot be trusted either, that model also needs to be evaluated. But if another model evaluates it again, then ultimately it becomes impossible to trust it. Understanding it from that perspective, I think the conclusion of model evaluation is that humans have to put in the hard work. - In fact, I understand that Meta hired more people and used more human labor to give tasks to and train models, and it seems that perspective is reflected here as well.

    • jhong
      Instructor

      Hello, thank you for leaving such a great review!! Since there are so many evaluation methods, I was worried that covering each and every one might feel like too high a barrier to entry, so I focused primarily on what they feel like and what direction they should take. So rather than focusing on the overall code structure, I think it would be great if you learned these patterns and approaches!! I’ll continue to provide more useful content in the future!!

  • kask814587762님의 프로필 이미지
    kask814587762

    Reviews 6

    ∙

    Average Rating 5.0

    5

    83% enrolled

    I found the lecture very informative!

    • achoi68339754님의 프로필 이미지
      achoi68339754

      Reviews 2

      ∙

      Average Rating 5.0

      5

      87% enrolled

      Summary: There are various ways to evaluate AI Agents, and all of them can be applied, but ultimately, evaluation still requires a lot of human involvement and judgment. I really enjoyed this take on such a fresh topic!!

      • leenux2890님의 프로필 이미지
        leenux2890

        Reviews 1

        ∙

        Average Rating 4.0

        4

        100% enrolled

        Hong's other courses

        Check out other courses by the instructor!

        Similar courses

        Explore other courses in the same field!

        Limited time deal ends in 6 days

        $34.10

        59%

        $84.70