Homework 1: Docker, SQL and Terraform
- Submissions
- 1578
- Completion
- 48.4%
- Median score
- 7
- Lecture time
- 6h
- Homework time
- 3h
- Total time
- 10h
Data Engineering Zoomcamp 2026
Course-level participation, homework, and project statistics.
Median time and score by homework. Hover values on desktop for interquartile ranges.
| Homework | Submissions | Completion | Lecture time | Homework time | Total time | Median score |
|---|---|---|---|---|---|---|
| Homework 1: Docker, SQL and Terraform | 1578 | 48.4% | 6h | 3h | 10h | 7 |
| Homework 2: Workflow Orchestration | 1021 | 31.3% | 5h | 3h | 9h | 6 |
| Homework 3: Data Warehousing | 873 | 26.8% | 3h | 3h | 7h | 8 |
| Homework 4: Analytics Engineering | 699 | 21.4% | 6h | 5h | 12h | 6 |
| Homework 5: Data Platforms | 613 | 18.8% | 4h | 2h | 7h | 7 |
| Workshop 1: Ingestion with dlt | 447 | 13.7% | 3h | 2h | 5h | 3 |
| Homework 6: Batch | 524 | 16.1% | 6h | 2h | 9h | 6 |
| Homework 7: Streaming | 426 | 13.1% | 5h | 5h | 11h | 6 |
Ranked by lowest median question score relative to the maximum achievable. Normalizing by question count makes homeworks of different lengths comparable; a lower percentage indicates a harder assignment.
| Rank | Homework | Median score | Score % | Completion |
|---|---|---|---|---|
| 1 | Homework 4: Analytics Engineering | 5 / 6 | 83.3% | 21.4% |
| 2 | Homework 5: Data Platforms | 6 / 7 | 85.7% | 18.8% |
| 3 | Homework 1: Docker, SQL and Terraform | 7 / 7 | 100.0% | 48.4% |
| 4 | Homework 2: Workflow Orchestration | 6 / 6 | 100.0% | 31.3% |
| 5 | Homework 3: Data Warehousing | 8 / 8 | 100.0% | 26.8% |
| 6 | Homework 6: Batch | 6 / 6 | 100.0% | 16.1% |
| 7 | Workshop 1: Ingestion with dlt | 3 / 3 | 100.0% | 13.7% |
| 8 | Homework 7: Streaming | 6 / 6 | 100.0% | 13.1% |
Share of submitted answers that were correct, per question. A lower percentage points to a harder or more confusing question. Participation-only questions are excluded.
| Homework | Question | Correct | Answers | % correct |
|---|---|---|---|---|
| Homework 1: Docker, SQL and Terraform | What's the version of pip in the python:3.13 image? | 1394 | 1578 | 88.3% |
| Given the docker-compose.yaml, what is the hostname and port that pgadmin should use to connect to the postgres database? | 1305 | 1578 | 82.7% | |
| For the trips in November 2025, how many trips had a trip_distance of less than or equal to 1 mile? | 1370 | 1578 | 86.8% | |
| Which was the pick up day with the longest trip distance? Only consider trips with trip_distance less than 100 miles. | 1245 | 1578 | 78.9% | |
| Which was the pickup zone with the largest total_amount (sum of all trips) on November 18th, 2025? | 1403 | 1578 | 88.9% | |
| For the passengers picked up in the zone named "East Harlem North" in November 2025, which was the drop off zone that had the largest tip? | 1171 | 1578 | 74.2% | |
| Which of the following sequences describes the Terraform workflow for: 1) Downloading plugins and setting up backend, 2) Generating and executing changes, 3) Removing all resources? | 1413 | 1578 | 89.5% | |
| Homework 2: Workflow Orchestration | Within the execution for Yellow Taxi data for the year 2020 and month 12: what is the uncompressed file size (i.e. the output file yellow_tripdata_2020-12.csv of the extract task)? | 598 | 1021 | 58.6% |
| What is the rendered value of the variable file when the inputs taxi is set to green, year is set to 2020, and month is set to 04 during execution? | 958 | 1021 | 93.8% | |
| How many rows are there for the Yellow Taxi data for all CSV files in the year 2020? | 937 | 1021 | 91.8% | |
| How many rows are there for the Green Taxi data for all CSV files in the year 2020? | 935 | 1021 | 91.6% | |
| How many rows are there for the Yellow Taxi data for the March 2021 CSV file? | 925 | 1021 | 90.6% | |
| How would you configure the timezone to New York in a Schedule trigger? | 941 | 1021 | 92.2% | |
| Homework 3: Data Warehousing | What is count of records for the 2024 Yellow Taxi Data? | 835 | 872 | 95.8% |
| What is the estimated amount of data that will be read when this query is executed on the External Table and the Table? | 739 | 872 | 84.7% | |
| Why are the estimated number of Bytes different? | 848 | 872 | 97.2% | |
| How many records have a fare_amount of 0? | 794 | 872 | 91.1% | |
| What is the best strategy to make an optimized table in Big Query if your query will always filter based on tpep_dropoff_datetime and order the results by VendorID (Create a new table with this strategy) | 835 | 872 | 95.8% | |
| Write a query to retrieve the distinct VendorIDs between tpep_dropoff_datetime 2024-03-01 and 2024-03-15 (inclusive). Use the materialized table you created earlier in your from clause and note the estimated bytes. Now change the table in the from clause to the partitioned table you created for question 5 and note the estimated bytes processed. What are these values? | 845 | 872 | 96.9% | |
| Where is the data stored in the External Table you created? | 834 | 872 | 95.6% | |
| It is best practice in Big Query to always cluster your data: | 809 | 872 | 92.8% | |
| Write a `SELECT count(*)` query FROM the materialized table you created. How many bytes does it estimate will be read? Why? | 0 | 872 | 0.0% | |
| Homework 4: Analytics Engineering | Q1: dbt run --select int_trips_unioned builds which models? | 491 | 699 | 70.2% |
| Q2: New value 6 appears in payment_type. What happens on dbt test? | 665 | 699 | 95.1% | |
| Q3: Count of records in fct_monthly_zone_revenue? | 478 | 699 | 68.4% | |
| Q4: Zone with highest revenue for Green taxis in 2020? | 611 | 699 | 87.4% | |
| Q5: Total trips for Green taxis in October 2019? | 559 | 699 | 80.0% | |
| Q6: Count of records in stg_fhv_tripdata (filter dispatching_base_num IS NULL)? | 577 | 699 | 82.5% | |
| Homework 5: Data Platforms | Bruin Pipeline StructureIn a Bruin project, what are the required files/directories? | 492 | 613 | 80.3% |
| Materialization Strategies You're building a pipeline that processes NYC taxi data organized by month based on pickup_datetime. Which incremental strategy is best for processing a specific interval period by deleting and inserting data for that time period? | 578 | 613 | 94.3% | |
| Pipeline VariablesYou have a variable defined in pipeline.yml:variables: taxi_types: type: array items: type: string default: ["yellow", "green"]How do you override this when running the pipeline to only process yellow taxis? | 582 | 613 | 94.9% | |
| Running with DependenciesYou've modified the ingestion/trips.py asset and want to run it plus all downstream assets. Which command should you use? | 381 | 613 | 62.2% | |
| Quality Checks. You want to ensure the pickup_datetime column in your trips table never has NULL values. Which quality check should you add to your asset definition? | 594 | 613 | 96.9% | |
| Lineage and DependenciesAfter building your pipeline, you want to visualize the dependency graph between assets. Which Bruin command should you use? | 529 | 613 | 86.3% | |
| Question 7. First-Time RunYou're running a Bruin pipeline for the first time on a new DuckDB database. What flag should you use to ensure tables are created from scratch? | 576 | 613 | 94.0% | |
| Workshop 1: Ingestion with dlt | What is the start date and end date of the dataset? | 381 | 447 | 85.2% |
| What proportion of trips are paid with credit card? | 382 | 447 | 85.5% | |
| What is the total amount of money generated in tips? | 363 | 447 | 81.2% | |
| Homework 6: Batch | Yellow November 2025 | 466 | 524 | 88.9% |
| Count records | 484 | 524 | 92.4% | |
| Longest trip | 462 | 524 | 88.2% | |
| User Interface | 499 | 524 | 95.2% | |
| Least frequent pickup location zone | 498 | 524 | 95.0% | |
| Homework 7: Streaming | Consumer - trip distance | 362 | 426 | 85.0% |
| Tumbling window - pickup location | 375 | 426 | 88.0% | |
| Session window - longest streak | 280 | 426 | 65.7% | |
| Tumbling window - largest tip | 329 | 426 | 77.2% |
When homework submissions arrive relative to the deadline, across all assignments.
Homework submissions per week, across all assignments. Bars are scaled to the busiest week.
581 students completed the course and received certificates.