Week 2 实战:基于 FastAPI + SQLite + Ollama 的 LLM 动作项抽取器(modern-software-dev-assignments 课程实践)
发布时间:2026/9/30 6:02:35来源:尧图网络
示例工程【免费下载链接】modern-software-dev-assignmentsAssignments for CS146S: The Modern Software Dev (Stanford University Fall 2026/2025)项目地址https://gitcode.com/GitHub_Trending/mo/modern-software-dev-assignments点击查看免费下载导读本文以 Stanford CS146S《现代软件开发》课程 Week 2 作业Action Item Extractor为骨架完整拆解如何在一个最小化的 FastAPI SQLite 应用中用 AICursor Ollama将自由文本笔记自动提取为可勾选的 action items。读者将从零走通启动现有应用 → 用 Ollama 实现 LLM 抽取 → 编写单元测试 → 重构后端 → 新增端点与前端按钮 → 自动生成 README → 填写 writeup 并提交的完整闭环并掌握结构化输出JSON Schema、pydantic 模型、数据库层清理与测试驱动开发的实战写法。文中的每一步都给出仓库源码路径与可直接复制的代码方便在本地逐行对照。1. 项目背景与当前应用概览1.1 这是什么项目本项目是 Stanford CS146S: The Modern Software DevFall 2026/2025的课程作业仓库README.mdWeek 2 的目标是在一个最小化的FastAPI SQLite应用之上通过 AI 辅助开发完成 5 个练习。仓库根目录下 pyproject.toml 声明了核心依赖依赖版本要求用途fastapi0.111.0Web 框架提供 REST 接口uvicornstandard0.23.0ASGI 服务器负责启动应用sqlalchemy / pydantic2.0.0ORM 与数据校验后续重构时使用python-dotenv1.0.0读取.env配置ollama^0.5.3调用本地 Ollama LLM 进行抽取openai1.0.0兼容 OpenAI 客户端协议备用开发依赖还包括 pytest、httpx、black、ruff、pre-commit 等为测试与代码质量把关。1.2 当前应用如何启动按 week2/assignment.md 的 Getting Started 步骤激活 conda 环境conda activate cs146s在仓库根目录运行服务器poetry run uvicorn week2.app.main:app --reload--reload让 uvicorn 监听源码变化自动重启适合开发迭代。浏览器打开 http://127.0.0.1:8000/会看到 week2/frontend/index.html 渲染出的前端页面一个文本框 Extract 按钮。输入笔记并点击 Extract观察预定义启发式规则如何抽取动作项。1.3 应用内部结构week2/app/main.py 是入口关键代码from .db import init_db from .routers import action_items, notes init_db() app FastAPI(titleAction Item Extractor) app.include_router(notes.router) app.include_router(action_items.router) app.mount(/static, StaticFiles(directorystr(static_dir)), namestatic)init_db()在应用启动时创建 SQLite 表notes与action_items。两个 APIRouter 分别挂在/notes与/action-items前缀下。前端静态文件挂在/static下/直接返回 week2/frontend/index.html 的内容。2. TODO 1Scaffold a New Feature —— 用 Ollama 实现extract_action_items_llm()2.1 先读懂现有启发式抽取当前核心逻辑位于 week2/app/services/extract.pyBULLET_PREFIX_PATTERN re.compile(r^\s*([-*•]|\d\.)\s) KEYWORD_PREFIXES (todo:, action:, next:)extract_action_items(text)按行扫描满足以下任一条件即视为动作行以-、*、•或1.等编号/项目符号开头以小写化后的todo:、action:、next:前缀开头包含[ ]或[todo]复选框标记。随后会剥掉前缀与复选框标记并在无匹配时回退为按句子切分、挑选以add/create/implement/fix/update...等祈使动词开头的句子最后做大小写不敏感的去重extract.py。这种方式对列表式笔记有效但对自然语言段落鲁棒性差——这正是 TODO 1 引入 LLM 的原因。2.2 设计思路与准备TODO 1 要求实现一个LLM 驱动的替代函数extract_action_items_llm()利用Ollama进行抽取。官方建议参考两份文档结构化输出JSON array of stringshttps://ollama.com/blog/structured-outputs —— 让模型严格返回 JSON 数组便于程序解析。模型库https://ollama.com/library —— 大模型更吃资源建议从小的模型开始。拉取并运行ollama run {MODEL_NAME}。项目依赖中已包含ollama ^0.5.3pyproject.toml可直接使用其 Python 客户端。2.3 可落地的实现骨架在 week2/app/services/extract.py 中追加注意保留原有extract_action_items供启发式路径使用并在代码注释中标明哪些部分是 AI 生成的这是评分要求之一# LLM-powered extraction (generated with Cursor) from ollama import chat from dotenv import load_dotenv load_dotenv() LLM_MODEL os.getenv(OLLAMA_MODEL, llama3.2) SYSTEM_PROMPT ( You are a note-processing assistant. Extract actionable items from the given free-form notes. Return ONLY a JSON array of strings, e.g. [Set up database, Write tests]. If there are no action items, return []. ) def extract_action_items_llm(text: str) - List[str]: response chat( modelLLM_MODEL, messages[ {role: system, content: SYSTEM_PROMPT}, {role: user, content: text}, ], formatjson, # structured output: force valid JSON ) content response[message][content] try: items json.loads(content) except json.JSONDecodeError: # Fallback: parse the first JSON array found in the response match re.search(r\[.*\], content, re.DOTALL) items json.loads(match.group(0)) if match else [] # Normalize and deduplicate like the heuristic path cleaned: List[str] [] seen: set[str] set() for item in items if isinstance(items, list) else []: if not isinstance(item, str): continue s item.strip() key s.lower() if s and key not in seen: seen.add(key) cleaned.append(s) return cleaned要点说明formatjson对应 Ollama 的结构化输出能力可显著降低解析失败的几率保留json.JSONDecodeError兜底用正则\[.*\]提取响应中的第一个 JSON 数组与启发式版本保持一致的去重策略保证两个函数的输出契约一致用os.getenv(OLLAMA_MODEL, llama3.2)让模型名可通过环境变量配置与项目已加载的load_dotenv()配合。2.4 运行验证确保本地已安装 Ollama 并拉取了模型然后ollama run llama3.2接着在 Python 中快速冒烟from week2.app.services.extract import extract_action_items_llm print(extract_action_items_llm(Meeting: - [ ] Set up db\n- Implement extract endpoint\nNext: write tests))预期得到类似[Set up db, Implement extract endpoint, write tests]的 JSON 数组。若 Ollama 未启动ollama.chat会抛出连接异常——这也是 TODO 3 中需要处理错误的动机之一。3. TODO 2Add Unit Tests —— 为extract_action_items_llm()补测试作业要求在 week2/tests/test_extract.py 中覆盖多种输入bullet lists、keyword-prefixed lines、empty input等。现有测试文件已经有一个针对启发式路径的用例test_extract_bullets_and_checkboxes我们需要在此基础上为 LLM 版本新增用例。注意直接调用真实 Ollama 会让测试依赖外部服务且速度慢。实践中建议使用monkeypatch或依赖注入把chat替换为假实现从而让测试快速、确定、可离线运行。示例# tests/test_extract.py (append) import json from ..app.services.extract import extract_action_items_llm def _fake_chat(model, messages, formatNone): class _R: def __getitem__(self, k): return {message: {content: json.dumps( [Set up database, Write tests])}}[k] return {message: {content: json.dumps( [Set up database, Write tests])}} def test_llm_extract_bullets(monkeypatch): import week2.app.services.extract as extract_mod monkeypatch.setattr(extract_mod, chat, _fake_chat) items extract_action_items_llm(Notes:\n- [ ] Set up database\n* Write tests\n) assert items [Set up database, Write tests] def test_llm_extract_keyword_prefixed(monkeypatch): import week2.app.services.extract as extract_mod monkeypatch.setattr(extract_mod, chat, _fake_chat) items extract_action_items_llm(todo: fix login\nnext: add logout) assert fix login in items def test_llm_extract_empty_input(monkeypatch): import week2.app.services.extract as extract_mod monkeypatch.setattr(extract_mod, chat, lambda **kw: {message: {content: []}}) assert extract_action_items_llm() []跑测试poetry run pytest week2/tests -v通过monkeypatch.setattr把chat替换为固定返回既验证了我们的解析/去重逻辑又避免了真实模型的不确定性——这也是为外部依赖打桩的通用测试模式。4. TODO 3Refactor Existing Code for Clarity —— 后端分层重构TODO 3 要求对整个后端做一次清晰度重构重点包括well-defined API contracts/schemas引入 pydantic 模型替代裸Dict[str, Any]database layer cleanup封装连接与 CRUD消除重复app lifecycle/configuration把init_db()从模块导入时执行改为 FastAPI lifespanerror handling统一 400/404 语义与异常处理。仓库中 week4/week5 等目录已经展示了重构后的形态week4/backend/app/schemas.py、week4/backend/app/models.py可对照参考最终目标。4.1 Schemas用 pydantic 定义 API 契约新建 week2/app/schemas.py以 week4 版本为蓝本from pydantic import BaseModel from typing import Optional class NoteCreate(BaseModel): content: str class NoteOut(BaseModel): id: int content: str created_at: str class ActionItemCreate(BaseModel): text: str note_id: Optional[int] None class ActionItemOut(BaseModel): id: int note_id: Optional[int] text: str done: bool created_at: str class ExtractRequest(BaseModel): text: str save_note: bool False class ExtractResponse(BaseModel): note_id: Optional[int] items: list[ActionItemOut]router 里把payload: Dict[str, Any]换成payload: ExtractRequest等模型后FastAPI 会自动完成校验、生成 OpenAPI schema并在缺字段时返回 422——这就是well-defined API contract。4.2 Database layer单一连接工厂与类型化返回现有 week2/app/db.py 已经提供了get_connection()row_factory sqlite3.Row、insert_note、list_notes、get_note、insert_action_items、list_action_items、mark_action_item_done等函数。重构方向将DB_PATH、DATA_DIR收敛到统一配置处或读.env所有函数保持函数内打开连接、with 块自动提交的既有模式避免长连接对list_notes/list_action_items返回的sqlite3.Row显式转换为 pydantic 模型让服务层不再直接接触数据库行对象。4.3 App lifecycle用 lifespan 替代模块级副作用当前 main.py 在模块导入时执行init_db()属于隐式副作用。重构后from contextlib import asynccontextmanager from .db import init_db asynccontextmanager async def lifespan(app: FastAPI): init_db() yield app FastAPI(titleAction Item Extractor, lifespanlifespan)这样数据库初始化只在应用启动时发生测试时也更容易控制。4.4 Error handling 统一将content 为空则 400这类校验逻辑交给 pydantic 的Field(min_length1)由框架返回 422对note not found保留HTTPException(404)对 LLM 抽取失败Ollama 未启动、JSON 解析失败在 service 层捕获并抛出可读的HTTPException(503, LLM service unavailable)避免 500 裸奔。5. TODO 4Use Agentic Mode to Automate Small Tasks —— 新端点 前端按钮TODO 4 分两步核心是让 AgentCursor Agentic Mode自动完成端到端接线5.1 新增 LLM 抽取端点在 week2/app/routers/action_items.py 中增加from ..services.extract import extract_action_items_llm router.post(/extract/llm) def extract_llm(payload: ExtractRequest) - ExtractResponse: text payload.text.strip() if not text: raise HTTPException(status_code400, detailtext is required) note_id db.insert_note(text) if payload.save_note else None items extract_action_items_llm(text) ids db.insert_action_items(items, note_idnote_id) return ExtractResponse( note_idnote_id, items[ActionItemOut( idi, note_idnote_id, textt, doneFalse, created_at ) for i, t in zip(ids, items)], )5.2 暴露检索所有笔记端点week2/app/routers/notes.py 目前只有POST /notes与GET /notes/{note_id}缺少列出全部笔记。新增router.get() def list_all_notes() - list[NoteOut]: rows db.list_notes() return [NoteOut(idr[id], contentr[content], created_atr[created_at]) for r in rows]5.3 前端加两个按钮在 week2/frontend/index.html 的.row中追加按钮并复制现有Extract的事件处理模式button idextract-llmExtract LLM/button button idlist-notesList Notes/button对应脚本$(#extract-llm).addEventListener(click, async () { const res await fetch(/action-items/extract/llm, { method: POST, headers: { Content-Type: application/json }, body: JSON.stringify({ text: $(#text).value, save_note: $(#save_note).checked }), }); // ...与 Extract 相同的渲染/勾选逻辑 }); $(#list-notes).addEventListener(click, async () { const res await fetch(/notes); const notes await res.json(); itemsEl.innerHTML notes.map(n div classitemspan#${n.id} · ${n.content}/span/div).join(); });注意现有前端用fetch(/action-items/extract, ...)和fetch(/action-items/{id}/done, ...)与后端直连index.html新按钮沿用同一套 fetch 风格即可。6. TODO 5Generate a README from the Codebase —— 让 AI 自动产出项目文档Learning Goal让学生体验 AI 如何剖析代码库并自动生成文档展示 Cursor 解析代码上下文、翻译成人可读形式的能力。在 Cursor 中对当前代码库发起类似下面的 Prompt让它生成 week2/README.mdAnalyze this FastAPI SQLite action item extractor codebase. Generate a README.md that includes: (1) a brief overview of the project; (2) how to set up and run the project; (3) the API endpoints and their functionality; (4) instructions for running the test suite. Use the actual code in week2/app, week2/tests and week2/frontend as evidence, and keep it concise.README 至少应覆盖对照作业要求逐条检查要求应包含内容Brief overview应用做什么把自由笔记转换为枚举的动作项清单Set up runconda activate cs146s、poetry install、poetry run uvicorn week2.app.main:app --reload、访问 8000 端口API endpointsPOST /notes、GET /notes、GET /notes/{id}、POST /action-items/extract、POST /action-items/extract/llm、GET /action-items、POST /action-items/{id}/doneTest suitepoetry run pytest week2/tests -v7. 填写 writeup 与提交按 week2/writeup.md 的模板逐项填写SUBMISSION DETAILSName、SUNet ID、Citations、耗时估算每个 Exercise 都要给出使用的 Prompt放进代码块与改动位置清单文件 行号尽量详尽尤其是 TODO 3 的散落改动在代码注释中明确标注哪些片段是 Cursor/AI 生成的。完成后自检Command (⌘) F或Ctrl F搜索文件中是否还有残留的TODO无结果即完成将所有改动 push 到远程仓库在 Gradescope 上提交。评分标准满分 100每个 Part 1-5 各 20 分生成代码 10 分 每个 Prompt 10 分——Prompt 质量和 writeup 完整度与代码同等重要。8. 常见坑与提示Ollama 未启动ollama run llama3.2会阻塞式启动若只拉模型用ollama pull llama3.2。调用chat()前确保服务在跑默认 http://localhost:11434。结构化输出兼容性formatjson需要较新的 Ollama 版本老版本可去掉format参数并依赖正则兜底解析。模型过大导致响应慢作业建议start small例如llama3.23B 参数档比 70B 档更适合单元测试。测试不要依赖真实网络用monkeypatch打桩chat见第 3 节。重构时保持行为不变先跑一遍现有 pytest 作为回归基线再逐层替换为 pydantic schema / lifespan。赞分享示例工程【免费下载链接】modern-software-dev-assignmentsAssignments for CS146S: The Modern Software Dev (Stanford University Fall 2026/2025)项目地址https://gitcode.com/GitHub_Trending/mo/modern-software-dev-assignments点击查看免费下载相关推荐如何用Ollama结构化输出构建LLM抽取端点modern-software-dev-assignments week2实战如何用Ollama结构化输出构建LLM抽取端点modern software dev assignments week2实战 本文带你用 Ollama 结构化示例工程用正则解析LLM输出modern-software-dev-assignments中的答案抽取工程实践用正则解析LLM输出modern software dev assignments中的答案抽取工程实践 modern software dev assignm示例工程modern-software-dev-assignments Week 5 起手项目用 FastAPI SQLite 静态前端搭建 Agent 驱动的极简全栈实验场modern software dev assignments Week 5 起手项目用 FastAPI SQLite 静态前端搭建 Agent 驱动示例工程创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考
网站建设高端定制企业官网