AI Engineering Buildcamp: from RAG to Agents Cohort 3

Homework 6: Evaluations Statistics

Distribution of scores and reported study time for this homework.

Submissions

7

Median total score

8

Average total score

7

Score distribution

All values are points.

Questions score

Min
6
Median
8.0
Max
8
Q1
7.0
Avg
7.4
Q3
8.0

Learning in public score

Min
-
Median
0.0
Max
-
Q1
0.0
Avg
0.0
Q3
0.0

Total score

Min
6
Median
8.0
Max
8
Q1
7.0
Avg
7.4
Q3
8.0

Time distribution

All values are hours reported by students.

Lectures

Min
-
Median
-
Max
-
Q1
-
Avg
-
Q3
-

Homework

Min
1.0
Median
2.0
Max
5.0
Q1
1.5
Avg
2.7
Q3
3.5

Question breakdown

Correctness and answer distribution per question.

1. Did the agent hallucinate on the almond milk substitution question?

6 / 7 correct (85.7%)

1 Yes, the agent made up substitution advice not in the recipe data 6 (85.7%)
2 No, the agent said it doesn't have substitution information 1 (14.3%)

2. Approximate cost of the almond milk substitution scenario

6 / 7 correct (85.7%)

1 Less than $0.001 6 (85.7%)
2 $0.001 - $0.01 1 (14.3%)
3 $0.01 - $0.05 0 (0.0%)
4 More than $0.05 0 (0.0%)

3. How many scenarios did the agent get wrong?

7 / 7 correct (100.0%)

Answer Count
2 3
12 1
14/36 1
9 1
10/24 bad. 1

4. How many results does the judge label as bad?

7 / 7 correct (100.0%)

Answer Count
20 1
7 1
26/36 1
The judge labeled 20 out of 30 as bad 1
19 1
14/24 bad. 1
0 1

5. Did you and the judge agree on the almond milk scenario?

5 / 7 correct (71.4%)

1 Yes, we both said good 1 (14.3%)
2 Yes, we both said bad 5 (71.4%)
3 No, I said good but the judge said bad 1 (14.3%)
4 No, I said bad but the judge said good 0 (0.0%)

6. Judge's reasoning for the almond milk scenario

7 / 7 correct (100.0%)

Answer Count
The assistant provided general cooking advice about substituting almond milk for regular milk in pancake recipes, which is beyond the provided recipe collection and is considered hallucination. It should have directly addressed the lack of a specific pancake recipe instead. 1
The assistant provided a general substitution suggestion that does not exist in the recipe collection and declined to answer about any available pancake recipes. This constitutes giving incorrect advice and the assistant should not have ventured beyond the recipe data. 1
The assistant provided general cooking knowledge about substituting almond milk in pancakes instead of referencing a specific recipe from the collection. It also offered advice on flavor balancing which is not derived from the recipe data. This constitutes a hallucination. 1
The assistant provided information about almond milk, which is not found in the recipe collection and therefore constitutes hallucination. It also suggested that almond milk can be used as a substitute in pancake recipes, which is cooking advice that isn't based on the recipe data. 1
The agent response incorrectly provided general cooking advice about substituting almond milk in pancakes, which is outside the scope of the recipe collection. It should not have suggested that almond milk can be used in most pancake recipes without issues, as this information is not found in the recipe data. 1
The agent correctly states that the recipe does not specify if almond milk can be substituted, indicating it is not making up information or giving advice outside of the provided recipe data. This is a good response because it accurately reflects the limitations of the recipe information without hallucinating or providing unverified cooking advice. 1
The response includes general cooking advice about using almond milk as a substitute, which is not information found in the recipe collection, constituting hallucination. Additionally, it incorrectly states that a pancake recipe is not available without checking the recipe data. 1

7. Number of disagreements before and after improving the judge

7 / 7 correct (100.0%)

Answer Count
before: 8, after: 3 1
2 1
12 before improvement of judge, 5 right after improvement. After improving, the difference are when agent tag as good for those items where agent is being extra helpful if no answer is found. Agent correctly says not available but tends to suggest workarounds. 1
6 1
19 before, 8 after 1
Before : 2 , After : 1 1
Disagreements before: 6 Disagreements after: 3 1

8. Percentage of bad labels from synthetic dataset (bonus)

7 / 7 correct (100.0%)

Answer Count
Around 59% 1
leave blank unless you run the synthetic pipeline 1
10 1
55/75 = 73.3% 1

Calculated: 11 June 2026, 20:15