Questions score
- Min
- 6
- Median
- 6.0
- Max
- 6
- Q1
- 6.0
- Avg
- 6.0
- Q3
- 6.0
AI Engineering Buildcamp: from RAG to Agents Cohort 3
Distribution of scores and reported study time for this homework.
Submissions
18
Median total score
6
Average total score
6
All values are points.
All values are hours reported by students.
Correctness and answer distribution per question.
18 / 18 correct (100.0%)
| Answer | Count |
|---|---|
| Project Title: AI-Powered Investor Intelligence Dashboard for PSX Overview: The project aims to build an automated investor decision-support dashboard that continuously scrapes and analyzes data from the Pakistan Stock Exchange (PSX), including: Company announcements Financial results (quarterly & annual reports) Corporate disclosures Market data The system will use AI + rule-based intelligence to convert raw financial and textual data into actionable investment signals. How It Works: 1. Data Ingestion Scrape PSX company pages (announcements, financials) Extract PDFs (using OCR + NLP) 2. Feature Extraction Financial metrics: EPS growth Free Cash Flow Debt levels Text signals: M&A announcements Expansion plans Board changes Regulatory risks 3. AI Layer RAG-based system to interpret reports LLM to: Detect turnaround signals Identify red flags Extract catalysts 4. Scoring Engine Each company is scored on: Fundamentals Cash flow strength Catalyst strength Risk signals 5. Dashboard Output Buy/Sell recommendation Confidence score Key reasons (explainable AI) | 1 |
| I will build an ATS Gap Analyser for job seekers to diagnose why their CV isn't getting responses and what to fix. Input: a CV (PDF upload) and a job description (text paste). Processing: an agent extracts required skills and keywords from the job description, scores the CV against them, identifies gaps, and generates a prioritised list of specific fixes plus an optimised cover letter. Output: a gap report with a match score, missing keywords, specific CV improvement suggestions, and a tailored cover letter. Success metric: LLM-as-judge scores the gap report quality and cover letter relevance against a ground truth dataset of 20-50 CV + job description pairs, targeting average quality score above 4/5. | 1 |
| I’m building an AI-assisted contract intelligence app for unions that turns collective bargaining agreements into structured, reviewable data. The app extracts time-sensitive rules—like grievance deadlines, escalation windows, arbitration timelines, contract expiration dates, negotiation notice periods, and reopener windows—so union officers can verify them and generate calendars, reminders, and grievance-specific timelines. The goal is to help unions avoid missed deadlines, track which contract rules apply over time, and prepare more strategically for negotiations without replacing human judgment. | 1 |
| The Problem: Autonomous AI agents frequently generate "hallucinated" or outdated code when tasked with quantum circuit design. Because Qiskit 1.x (released in early 2026) introduced massive breaking changes from legacy versions, standard AI models often suggest deprecated functions (like qiskit.execute) that cause runtime crashes. This "version drift" creates expensive errors and wasted classical/quantum compute resources. Who Has It: Enterprise AI Solutions Architects and Quantum Researchers who need a governance layer to "steward" autonomous agents, ensuring that generated code is safe, up-to-date, and compliant with the NIST AI 600-1 Agentic Profile before execution. A Typical Interaction What the User Provides: The user (a developer or an upstream AI agent) provides a high-level intent, such as: "Build a 3-qubit teleportation circuit and optimize it for the 'ibm_cleveland' backend using the latest primitives." What the System Returns: The Quantum Bridge returns a structured Stewardship Packet including: Verified Code: Executable Python code that has been autonomously "self-corrected" by the agent via a local Qiskit 1.x linter and simulator. The Governance Audit: A NIST-aligned "Decision Probe" log explaining why specific gates were chosen and confirming that the code avoids deprecated functions. Visual Scaffolding: A Mermaid-rendered circuit diagram so the human user can visually verify the entanglement and gate sequence before hitting "Run." | 1 |
| I'm planning to build a study partner AI. It will know what I'm learning and my goals, and actively test my retention — through quizzes, concept reviews, or Socratic questioning — rather than just letting me passively consume material. The problem: In the age of AI, it's easy to accumulate cognitive debt. It's also easy to fall into tutorial hell when learning on your own. Even with structured courses, we often finish them without truly retaining the information afterward. A typical interaction: I'm envisioning a CLI tool I can chat with, since most of what I'm learning is coding-related. Starting as a CLI tool, with potential to expand to a web interface. | 1 |
| AI Tracker — a personal AI news aggregator that pulls daily updates on LLMs, deep learning, and agentic AI from multiple public sources (Hacker News, Reddit, arXiv, GitHub trending, RSS feeds from major AI labs, X, Bluesky, Hugging Face Papers, Polymarket). It deduplicates, scores by relevance keywords, retains items for a configurable window, and produces a daily digest. Backend is Python/FastAPI with async SQLAlchemy; LLM-powered paper summaries and notebook generation use Kimi K2.5. The goal is to replace ad-hoc browsing across a dozen sites with one curated feed I actually read. | 1 |
| Problem: Planning a backpacking trip is overwhelming. Information is scattered across dozens of tabs, blogs, and Reddit threads, and most of it is written for tourists, not backpackers. Backpackers have specific needs: budget hostels, cheap local food, public transport, visa costs, and what to actually pack. There's no single place that gives it to them clearly. Who has it: Solo backpackers and long-term budget travelers who live out of a single bag and need practical, no-fluff information fast. Typical interaction: User provides: a destination (or a few they're considering), optionally their nationality, budget range, or trip length System returns: a concise backpacker-focused breakdown — daily cost estimates, transport options, hostel neighborhoods, food spots, visa requirements, and on-the-ground tips | 1 |
| I want to build Applied ML Teaching Copilot, an AI assistant that helps instructors and students work with the materials of an Applied Machine Learning course. The problem is that technical course content is usually distributed across slides, notebooks, readings, assignments, and instructor notes, which makes it difficult to quickly prepare lectures, answer student questions, or generate reliable study material grounded in the actual course content. A typical interaction would be: the user asks a question such as “When should I use MAE instead of MSE?” or requests an output such as “Create a short study guide about decision trees.” The system retrieves relevant passages from the indexed course materials and returns a clear explanation, examples, key points, and study-oriented guidance based on those materials. | 1 |
| **Problem & user** Students and early-stage policy learners struggle to understand EU climate policy because key information is spread across multiple complex and technical documents, such as the European Climate Law and related EU policy pages. This makes it difficult to extract accurate, connected, and trustworthy answers without misinterpretation. **Typical interaction** A user asks a question like: *“How does the EU’s 2030 climate target relate to the 2050 goal?”* The system uses an **Agentic RAG approach** to retrieve relevant passages from official documents, decides whether the information is sufficient, and returns a clear, concise answer with citations. If the documents do not contain enough information, the system either explicitly states that or supplements the answer using a web search tool with external sources. | 1 |
| A capstone project for the AI Engineering Buildcamp. This system automatically generates structured clinical radiology reports for knee osteoarthritis X-rays by combining a domain-adapted Vision-Language Model (MedGemma) with Multimodal Retrieval-Augmented Generation (RAG). The Problem Radiologists manually grade knee osteoarthritis severity using the Kellgren-Lawrence (KL) scale (0–4), a process that is time-consuming, subjective, and inconsistent across practitioners, yet it directly determines which patients qualify for surgery. It's an AI system that generates radiology reports for knee X-rays. ( I can modify this project later) | 1 |
| I will build a WhatsApp-based data assistant that allows users to query a relational data warehouse using both natural language and reusable command shortcuts. Problem: - Business users cannot easily retrieve data from a complex data warehouse without relying on data engineers, due to the need for SQL and fragmented reporting tools. - Data engineers spend significant time repeatedly answering similar questions and manually querying data, leading to inefficiency. - There is no simple, trusted interface where users can quickly access validated data insights or reuse common queries. Input: User messages via WhatsApp, either as natural language queries or predefined commands (e.g., /today, /revenue). Processing: Classify user intent; retrieve relevant schema and similar past queries from a query memory system; generate and validate SQL using retrieval-augmented generation; execute queries in DuckDB; and compute trust signals based on similarity to previously validated queries. Output: Concise answers delivered in WhatsApp, including results, explanations, SQL traces, and trust indicators (e.g., similar queries, tables used). Success metric: Reduction in repeated manual queries and improved user trust, measured by reuse of shortcuts and accuracy on predefined business questions. | 1 |
| Personal finance agent that analyses your bank transactions and tells you where your money is going, where you can cut costs and suggests an investment portfolio based on your financial goals. An example is the Cleo app in the UK. - The agent asks you about your financial goals, how much you make, how much you spend etc. I’ll put a maximum of 3 questions, or maybe a checkbox is better. - You are then asked to connect your bank account. I’ll use the Plaid api get the bank transactions. - Then i feed the transactions data to Claude api to suggest some cost cutting or make a breakdown of where money is being spent - Based on this data, i’ll then suggest an investment portfolio | 1 |
| Ambitious self-learners and early-career professionals often save large amounts of useful content across platforms such as YouTube Watch Later, notes apps, bookmarks, and articles, but struggle to turn that information into structured progress. They face knowledge overload, unclear priorities, and inconsistent execution because their resources are scattered and there is no system that connects saved content to a clear learning roadmap. I would build an AI Learning Chief of Staff with a Knowledge Brain. The user provides their learning goal (e.g., become job-ready in AI engineering in 4 months), available study time, and access to saved resources such as YouTube links, notes, and articles. The system ingests and organizes the content into a searchable personal knowledge base, generates a roadmap, ranks the most relevant resources, and returns a weekly action plan with recommended next steps. A typical interaction would be: the user asks “What should I focus on this week for AI engineering?” and the system returns a personalized study plan, prioritized resources from their own saved content, and progress tracking updates. | 1 |
| ### Lucid AI Learning Reader *Read it. Understand it. Remember it.* This project is a personal AI-powered reading companion designed to make dense technical material — AI books, research papers, and whitepapers — genuinely accessible. Built with a dyslexia-first mindset, the tool meets you at the start of every chapter with a plain-English summary to prime your understanding before you read, then quizzes you at the end to reinforce retention. What makes it different from a generic Q&A chatbot is that every question, summary, and explanation is strictly grounded in the document you uploaded — the system is architecturally prevented from mixing in outside knowledge, so if you're reading about a specific software version or a paper's particular methodology, the quiz reflects exactly that, nothing more. The learning layer goes beyond one-off quizzes. The system tracks your quiz history across sessions and uses spaced repetition scheduling to resurface questions you struggled with at the right intervals — similar to how Anki works, but generated automatically from your own reading material. For whitepapers loaded with academic jargon, it auto-generates a glossary of technical terms simplified to an undergrad reading level, and you can ask follow-up questions about any concept at any time without losing your place in the structured reading flow. Core features: - Chapter pre-summary generated before you read, written in plain English to frame the content - End-of-chapter quiz with explanations on incorrect answers - Spaced repetition that persists across sessions and resurfaces weak areas automatically - Whitepaper simplifier that rewrites PhD-level language to undergrad level - Auto-generated glossary with the ability to ask follow-up questions on any term - Free-form Q&A available at any point during reading - Multi-layer hallucination prevention so every question is traceable back to the source text - Learning dashboard showing retention scores, session history, and progress by document | 1 |
| Project idea - GapFinder The Problem People who learn from long-form videos (e.g., students, self-learners, engineers watching tutorials) often feel like they understand the material but can’t identify what they’ve actually missed. This leads to inefficient rewatching and shallow learning because there’s no clear feedback on gaps in understanding. What It Does The user provides a YouTube video link and answers a small set of generated questions about the content. The system analyzes their responses against the video’s key concepts and returns a structured report highlighting what they understood well, what they misunderstood or missed, and which specific parts of the video they should revisit. What the system actually does Input: YouTube video URL User answers to questions System flow Step 1 — Extract & structure knowledge Transcribe video Break into concepts (chunking + labeling) Step 2 — Generate diagnostic questions Not generic questions — but: Concept coverage questions “Explain in your own words” prompts Application questions (transfer knowledge) Step 3 — User answers User types responses Step 4 — Gap detection (core innovation) The system compares: Expected concepts (from transcript) User answers And identifies: Missing concepts Misunderstandings Shallow explanations Agent structure: Planner Agent Decides: Which concepts to test Which question types to generate Question Generator Tool: Creates diagnostic questions Evaluation Tool (LLM-as-judge): Grades answers against concept checklist Gap Analyzer Tool: Maps errors → missing concepts | 1 |
| Problem Statement: Job seekers often struggle not only to find relevant job opportunities but also to understand how well their profile matches those roles and what they need to improve. Existing platforms like LinkedIn and Indeed help users discover jobs, but they do not provide personalised, actionable insights on skill gaps, profile alignment, or how to improve chances of getting shortlisted. This problem affects students, early-career professionals, and experienced candidates who apply to multiple roles without clear feedback on why they are not a good fit. Typical Interaction- The user provides: - Their CV (resume text or PDF) - Optional preferences (target role, job title, location, seniority level) The system: - Retrieves a list of relevant job postings based on the user’s profile and preferences - Extracts and structures key skills and experience from the CV - Analyses requirements from selected job descriptions - Compares the user’s profile with job requirements to generate: - A match score for each job - A clear gap analysis (missing or weak skills) - Actionable suggestions to improve the CV/profile - Optional interview preparation questions tailored to the role The output is a structured report that helps the user both discover suitable jobs and understand how to improve their chances of success for each role. | 1 |
| The problem: Heavy news consumers — people who follow AI, tech, markets, and sports across many sources — regularly read articles that mention future-dated events (court verdicts, product launches, earnings calls, sports roster announcements, movie releases) and then forget those dates by the time they arrive. There is no tool that extracts these dates, reminds them on the day, and verifies the date hasn't shifted before firing the reminder. The initial user is me, but the pattern fits any information-heavy knowledge worker who tracks unfolding stories. Typical interaction: The user forwards a news article (URL or text) to a Telegram bot. Within seconds, the bot replies with a confirmation listing the future-dated events it extracted from the article — for example, "I'll remind you about: (1) Adani SEBI verdict on May 15, 2026, (2) hearing on July 3, 2026." One to two days before each event, the system runs a freshness-check agent that searches for recent news on the event and decides whether the date has shifted, the event was cancelled, or it's still on track. If anything has changed, the user gets a proactive Telegram alert ("Heads up — the verdict has been postponed to June 2") and the reminder reschedules itself. On the event day, the bot sends the final reminder with a one-line summary and a link back to the original source. | 1 |
| I'm building an AI Diet Coach Agent for people trying to lose weight. The user provides their current weight, target weight, deadline, dietary restrictions, daily schedule, and location. The agent interviews them, then generates a personalized weekly meal plan — recipes for days they can cook, and nearby restaurant suggestions for busy days. When plans change, the agent re-plans automatically. | 1 |
18 / 18 correct (100.0%)
| Answer | Count |
|---|---|
| https://github.com/salman1127/ai-engineering-buildcamp-project | 1 |
| https://github.com/Amar-Ag/ats-gap-analyser | 1 |
| https://github.com/MikiYamFos/cba-clock | 1 |
| https://github.com/MsBluphire/quantum-bridge-capstone | 1 |
| https://github.com/dinobronx/ai-buildcamp-project-starter | 1 |
| https://github.com/hgiang/ai-tracker | 1 |
| https://github.com/whassan99/backpacker-ai | 1 |
| https://github.com/larsvasseldonk/relational-rag | 1 |
| https://github.com/wesleytanjiale/ai-learning-os/tree/main | 1 |
| https://github.com/davidacodes/Lucid-AI-Learning-Reader/ | 1 |
| https://github.com/CarSomma/ai-buildcamp-project-starter | 1 |
| https://github.com/saig217/future-event-remainder | 1 |
| https://github.com/55382/ai-buildcamp-project-starter | 1 |
| https://github.com/marcoteran/applied-ml-teaching-copilot | 1 |
| https://github.com/katjaweb/gapfinder | 1 |
| https://github.com/nehabarve0304/ai-buildcamp-capstone-project | 1 |
| https://github.com/thetsuwin66/ai-diet-coach-agent | 1 |
18 / 18 correct (100.0%)
| Answer | Count |
|---|---|
| Pakistan Stock Exchange Company profiles Announcements Financial reports | 1 |
| A hand-crafted ATS best practices knowledge base with 10 documents covering topics like keyword matching, CV formatting for ATS, tailoring CVs to job descriptions, cover letter best practices, and common rejection reasons. Documents are stored directly in the notebook as a list of dicts with title, category, and content fields. No external API needed — the knowledge base is static and curated. | 1 |
| The code for the cba_downloader for one website works and I have downloaded 10 cbas. They are hidden behind .ignore but here are the paths for a couple of them so you know they are real! data/samples/raw_pdfs/1012_DOI_BIA_Indian_Educators_Federation_4524_10042019-_redacted.pdf data/samples/raw_pdfs/ARS_Animal_Disease_Center2015-06-29BUS_1529.pdf | 1 |
| Synthetic Documents: a curated JSON file containing simulated Qiskit 1.x migration rules, NIST 600-1 compliance requirements, and circuit design patterns. query = "How should I execute my 3-qubit teleportation circuit on ibm_cleveland in Qiskit 1.x? Can I use qiskit.execute?" results = To execute your 3-qubit teleportation circuit on `ibm_cleveland` in Qiskit 1.x, you cannot use `qiskit.execute`, as it has been deprecated. Instead, you should use the V2 Primitives, specifically `qiskit.primitives.Sampler` for sampling circuits. To select the backend, utilize `QiskitRuntimeService` to retrieve the IBM Quantum backend. For example, you can get the `ibm_cleveland` backend using `service.backend('ibm_cleveland')`. Make sure your final payload does not include any deprecated functions to comply with NIST AI 600-1 guidelines. | 1 |
| Downloaded from swift documentation https://github.com/dinobronx/ai-buildcamp-project-starter/blob/main/data/swift_docs.json | 1 |
| I'm using Wikivoyage, the free community-written travel guide that's a sister project to Wikipedia, as my knowledge base. Wikivoyage city pages are already structured into the exact sections backpackers care about (Get in, Get around, Sleep, Eat, Buy, Stay safe, Stay healthy, Connect, Go next, etc.), which makes it a natural fit for section-level retrieval. How I get it: the notebook calls the public Wikivoyage MediaWiki API at https://en.wikivoyage.org/w/api.php with prop=extracts&explaintext=1&exsectionformat=wiki to pull plain-text page extracts. No API key, just a User-Agent header. I fetch ten popular backpacker hubs (Bangkok, Chiang Mai, Hanoi, Ho Chi Minh City, Lisbon, Barcelona, Mexico City, Bali, Kathmandu, Cusco) and cache each one to data/wikivoyage/<City>.json. The fetch cell is idempotent, so re-running it only hits the network for cities that aren't already on disk. A second cell splits each cached page on its == Section == headers, keeps a 17-section allowlist of backpacker-relevant sections, and produces 158 section-level documents of the form {city, country, section, title, content}. These are indexed with minsearch.Index(text_fields=["title", "section", "content"], keyword_fields=["city", "country"]), so each retrieval result is one focused chunk like "Chiang Mai / Get in" or "Lisbon / Sleep". Query: "I'm in Bangkok and want to get to Chiang Mai cheaply, what are my options and roughly what should I budget?" Answer the notebook produced: To get from Bangkok to Chiang Mai cheaply, you have a couple of budget-friendly options: By Bus: - Duration: approx. 9 to 12 hours, depending on the bus type. - Cost: 488 to 550 baht for first-class buses like Nakhonchai Air. Government buses are cheaper but less comfortable; buy tickets at Mo Chit Bus Terminal in Bangkok. - Tip: avoid buses advertised as "VIP" by travel agents on Khao San Road, as they might be inferior. By Train: - Duration: 12 to 15 hours; consider an overnight train to save on accommodation. - Cost: - Third-class: 231 baht - Second-class: 391 baht (non-AC) - Second-class with AC: 641 baht - Sleeper fares vary from 771 to 1653 baht, depending on class and position (top/bottom bunk). - Tip: book in advance, especially for the popular overnight sleeper trains. Additional Tips: - Arrival in Chiang Mai: you will likely arrive at Arcade Bus Station or Chiang Mai Train Station, both of which are not far from the city center. - Local Transport: use the minibus or songthaew (shared taxi) from the station to get to your accommodation. Budget: for bus or train travel plus local transport, budget roughly 500 to 1000 baht total, varying based on your choice and travel class. This information is based on the following sections: Chiang Mai (Get in) and Bangkok (Get in). | 1 |
| Digging into space repetition and dyslexia https://en.wikipedia.org/wiki/Spaced_repetition https://blog.dyslexia.com/dyslexia-and-comprehension-when-the-words-dont-stick/ I think "Attention is all you need" would be a great paper for the demo. https://proceedings.neurips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf | 1 |
| The system uses news article text (fetched on demand from URLs the user forwards via Telegram), structured event data extracted from those articles using Anthropic's Claude API with Pydantic schemas, fresh search results retrieved by the freshness-check agent via the Brave Search API and HTTP fetching with trafilatura, telemetry from Pydantic Logfire and a local SQLite database, and a hand-labeled evaluation dataset of 50–100 articles with ground-truth events. All external APIs (Anthropic, Telegram, Brave) operate within free or low-cost tiers; no scraping at scale, no proprietary datasets, no PII beyond the user's own Telegram chat ID. | 1 |
| https://www.kaggle.com/datasets/shashwatwork/knee-osteoarthritis-dataset-with-severity | 1 |
| Hacker News, Reddit, arXiv, GitHub trending, RSS, X (Twitter), Hugging Face Papers, LLM enrichment (paper summaries, notebook generation) calls the Kimi K2.5 API (api.moonshot.ai/v1) | 1 |
| For this first version, I used an initial synthetic course-material dataset for Applied Machine Learning. The dataset is stored in data/course_materials.json and contains short course-note records about topics such as MAE vs MSE, decision trees, Gini impurity, entropy, overfitting, train/test split, cross-validation, classification metrics, imbalanced classification, and RAG basics. I used synthetic data at this stage to quickly validate the end-to-end RAG pipeline with a controlled and well-structured prototype dataset. The data follows a simple format with fields such as id, module, lesson, topic, content, and source_type. This allows the notebook to search over topic and content while keeping useful metadata such as module, lesson, and source type. In the next iteration, I plan to replace and expand this synthetic dataset with real course materials from my Applied Machine Learning teaching repositories, especially: https://github.com/marcoteran/machinelearning https://github.com/marcoteran/ml https://github.com/marcoteran/iot These repositories will provide real class materials such as notebooks, slides, examples, exercises, and course notes. The goal is to progressively turn the system into a teaching copilot grounded in my actual Machine Learning course content. Example query: When should I use MAE instead of MSE in a regression problem? Notebook answer: You should use Mean Absolute Error (MAE) instead of Mean Squared Error (MSE) in a regression problem when you prefer an evaluation metric that is easy to interpret in the target units and when you want to be less sensitive to outliers. MAE calculates the average of the absolute errors, treating each error linearly, which makes it straightforward to understand the average magnitude of errors in the same units as your target variable. This is particularly useful when clarity in communication of the model's performance is essential. In contrast, MSE squares the errors, which means that larger errors disproportionately affect the overall score. Therefore, MSE is more appropriate when large errors are particularly costly, or when you want to penalize those large misses more heavily during model training and evaluation. In summary, choose MAE for interpretability and robustness against outliers, and choose MSE if you need to heavily penalize large errors. | 1 |
| The project uses YouTube video transcripts as its external data source, extracted via the YouTube API. This data is used to generate core concepts by extracting and structuring the key ideas from the transcript and for chunking for gap detection. | 1 |
| I generated a synthetic sales database (sales.duckdb) with three tables — customers, products, and orders — containing ~40 rows total, serving as a realistic stand-in for a business reporting database. Example query: "Which region had the highest revenue, and how does each region compare?" Generated SQL: SELECT region, ROUND(SUM(o.quantity * p.unit_price), 2) AS total_revenue FROM orders o JOIN products p USING (product_id) GROUP BY region ORDER BY total_revenue DESC Explanation (from the LLM): "This query calculates total revenue for each region and orders the results by revenue in descending order to identify the region with the highest revenue." | 1 |
| Fetch bank transactions data - Plaid api Giving suggestions on where to cost cut - feed the transactions to Claude api ask it to analyse and make suggestions. For investment portfolio suggestions - Not sure if i should use an api for this part or logic | 1 |
| Data chosen: 25 synthetic AI engineering learning resources, stored in data/resources.json. Each document represents a resource a self-learner might save — YouTube videos, articles, tutorials, and courses. Fields are title, type, topic, description, difficulty, and url. Topics covered: RAG, LLMs, prompt engineering, embeddings, vector databases, fine-tuning, agents, evaluation, MLOps, and Python. Synthetic was the right choice for a prototype because it's instantly runnable and already shaped like real user data. Answer the notebook produced: 1. Prompt Engineering Guide (2 hours) — OpenAI's guide covering clear instructions, task splitting, and systematic testing. Relevant as a beginner foundation before touching more complex topics. 2. Advanced Prompt Engineering: Chain of Thought and Self-Consistency (3 hours) — Deep dive into reasoning-based prompting techniques. Strengthens your ability to design prompts for complex AI applications. 3. Prompt Injection Attacks and Defenses (2 hours) — Simon Willison's analysis of prompt injection risks and the dual-LLM defense pattern. Essential reading before deploying any system that processes user input. 4. LLMOps: Monitoring and Observability in Production (3 hours) — MLflow Tracing overview for logging traces, tracking token costs, and debugging RAG failures in production. Total: 10 hours. Objective: solid foundation in prompt engineering and deployment practices before moving to RAG and agents in later weeks. | 1 |
| The project will use a combination of external job listing data and curated datasets to support job retrieval and gap analysis. For the proof of concept, I will start with a static dataset of job descriptions (e.g., JSON/CSV) containing structured fields such as job title, required skills, and experience level. This ensures consistent, high-quality data for developing and testing the gap analysis system. To simulate real-world usage, I may also integrate a free public job API such as Arbeitnow or Adzuna (via platforms like RapidAPI) to retrieve live job listings based on user preferences. Additionally, publicly available datasets (e.g., from Hugging Face) may be used to enrich job descriptions or provide standardised skill lists. This combination allows the system to balance reliability (static data for development and evaluation) with real-world relevance (API-based job retrieval). | 1 |
| I used the TheMealDB free API (https://www.themealdb.com/api.php) to source recipe data. I fetched 201 recipes across 14 food categories (Beef, Chicken, Seafood, Vegan, Vegetarian, etc.) and 6 Asian cuisines (Chinese, Japanese, Thai, Vietnamese, Malaysian, Filipino). Each recipe includes name, category, cuisine area, ingredients, cooking instructions, and an estimated cooking time in minutes. The data is stored in data/recipes.json and indexed with minsearch for the RAG pipeline. | 1 |
Calculated: 11 May 2026, 14:54