Translating frontier AI into products people can use. 把前沿 AI 能力,翻译成真实可用的产品体验
你好,我是袁泽华,浙江大学人工智能硕士在读,本科毕业于华中科技大学网络空间安全学院。
我关注 AI 产品、模型评测与 Agentic Workflow,也在持续思考如何让 Agent 真正进入产品、研发、设计和运营等岗位的日常工作流。
我会以每周更新的节奏,在这里整理项目、文章、前沿技术观察和阶段性思考。
参与生成式创作与 Agent 场景的评测归因,把主观体验、链路错误和模型边界拆解成可复用的产品判断。
用 Figma、Codex、MCP 与 Agent 工具链推进 Demo、SOP 和内部提效,让 AI 成为不同岗位可接入的工作流能力。
参与生成式创作产品 0-1 验证,负责需求拆解、模型与 Agent 评测、Demo 交付和团队 AI 提效沉淀。
AI Product · Internship学院综合成绩排名 3 / 176,关注 AI 产品、模型评测与医疗 AI。
AI · Product网络空间安全本科;曾获人民奖学金、明德奖学金、新生奖学金与优秀毕业生。
EducationWorked on FedPMR, a personalized prototype-based federated learning framework for multi-center accelerated MRI reconstruction. The system supports cross-institution model collaboration while keeping raw medical data local.
Result: Validated on three public datasets and one private dataset from The First Affiliated Hospital, Zhejiang University. Communication cost was reduced from 236.70 Mb to 8 Kb, with stronger reconstruction performance on heterogeneous sites.
Authors: Jing Wang, Tong Wang, Siyuan Luo, Xiaoran Guo, Jiayi Yin, Jing Li, Zehua Yuan, Gang Yu, Bo Lin.
提出一种面向临床多源异构数据的冠心病多模态风险预警方法。针对心电信号(ECG)、超声心动图视频与结构化临床表格数据三类异构模态,设计独立的深度特征提取网络,并通过 InfoNCE 对比学习实现跨模态特征空间对齐。
Innovation:引入基于可学习标识符的模态丢弃策略解决真实临床中的整模态缺失问题,结合交叉注意力机制进行特征动态加权融合,最终通过深度生存分析输出个体化动态生存曲线,为临床分诊提供时间维度的量化依据。
Field:数字医疗 · 医学人工智能 · 多模态深度学习
一个前端-only 的 Coding Agent 工作流 Demo:把 Goal、Task、Agent、Evaluation Criteria、失败处理与 Timeline 组织成可检查、可编辑、可导出的控制台。
Built a React + TypeScript MVP for managing an agentic coding loop from goal definition to task ownership, criteria evaluation, failure handling, replanning, and traceable export.
Product value: Makes evaluation standards explicit before and after implementation, helping AI-assisted development move from scattered prompts to an inspectable workflow with task-level acceptance criteria.
一个面向 AI 生成图像策略评估的在线盲测平台,用于把“看图打分、方案对比、结论导出”这类评审流程产品化。
Built a lightweight web-based blind testing platform for evaluating AI-generated images across multiple strategies. Reviewers can score anonymous options, add quick labels and comments, then reveal strategy identities after review.
Product value: Turns subjective visual comparison into structured, exportable evaluation data, making model and strategy discussions easier to align across product, design, and technical teams.
一个把个人主页虚拟形象 Loopi 做成可评测、可迭代 AI 产品实验的开源系统:真实访客反馈进入数据库,Agent 读取反馈、评估红线、生成候选并给出 staging 建议。
Built a personal-brand AI companion system for yuanzehua.me, turning a homepage mascot into a measurable product loop with visitor feedback, version records, candidate prompts, evaluation reports, and human-gated release decisions.
Product value: Demonstrates how AI-native product work can connect hypothesis, measurement, controlled generation, evaluation, and deployment without letting novelty override homepage fit or personal-brand clarity.
一个面向个人网站长期维护的 Codex Skill,把“读项目、找源文件、控制实验、验证改动、谨慎发布”沉淀成可复用的 AI 协作流程。
Built a reusable Codex skill for maintaining personal websites as long-lived personal products. It separates inspection, maintenance, experimentation, publication, and rollback, helping Codex avoid editing generated output or publishing without explicit authorization.
Workflow value: Turns repeated homepage updates into a source-aware, verification-first operating system, making AI collaboration safer for content updates, visual experiments, production checks, and rollback planning.
Built a knowledge-base Q&A system over PDF, DOCX, and Markdown documents using OpenAI APIs.
Design: Combined semantic chunking, embeddings, vector retrieval, and prompt design to reduce omissions and hallucinations in complex document Q&A, turning raw materials into an interactive knowledge base.
Designed a FedAvg-based multi-client training framework for privacy-sensitive behavior recognition, enabling model collaboration without moving local data.
Balance: Focused on the tradeoff between model performance, data locality, privacy protection, and deployability.
这是一段脱敏后的 AI 创作工具 0-1 案例:我负责把模糊的审美体验拆解成可评测指标,并用盲测数据支持产品判断。
同时,我用 AI 原生工具链把产品想法推进到可运行 Demo 和 SOP 沉淀,并探索 Agent 如何辅助产品、设计、研发和业务同学完成更短路径的协作。
Worked on a confidential 0-1 AI creative tool, covering user input, style control, color constraints, multi-turn editing, result evaluation, and creative workflow design.
Evaluation: Translated subjective aesthetics into measurable dimensions including color consistency, color hierarchy, and semantic fit. Led blind tests across 3 candidate model routes, collecting 343 ratings on 230 experimental images to support evidence-based product decisions.
AI-native delivery: Used Figma, Codex, MCP, and Vibe Coding workflows to independently build an interactive demo, then documented a 28-page SOP to help the team reuse Agent-assisted workflows for faster reviews, prototype validation, and cross-role collaboration.
Agent thinking: Paid close attention to how Agent systems fail across intent understanding, tool use, intermediate reasoning, output verification, and final presentation, turning evaluation from simple scoring into actionable workflow diagnosis.
Contributed to a smart greenhouse digitization project, from county-level rollout planning to technical implementation design.
Delivery: Integrated data flows from two greenhouses, a compact weather station, and the ULAND IoT enablement platform, turning environmental monitoring and weather alerts into planting decision support.
Supported front-desk service and information-system maintenance at a university counseling center. The experience shaped my understanding that good systems should be efficient, privacy-aware, and humane.
Loopi 是 yuanzehua.me 的动态品牌实验:访客反馈进入数据库,四个 Agent 定时读取、评估、生成候选并选择 staging 版本。现在可以直接观看这个 Loop 如何运转。
正在读取反馈统计...
Loop Canvas 是一个前端-only 的 Agent 工作流控制台:把 Goal、Task、Agent、Evaluation Criteria、失败处理与 Timeline 串成可检查、可编辑、可导出的协作界面。
关注 LLM/VLM、Agent、Text2SQL、多模态与医疗 AI。既想理解前沿能力如何涌现,也在意它们最终能为谁解决什么问题。技术深度是产品判断力的底座。
好的产品决策来自对用户、技术与商业三者之间张力的持续感知。我尤其在意如何把非标准化场景、Agent 链路和主观体验变成可讨论、可量化、可验证的判断体系。
AI 能力需要好的界面和工作流才能被稳定使用。关注提示词、风格控制、多轮编辑、工具调用、反馈与置信度表达,让智能变得可控、可信赖,而不是神秘。
我也喜欢把复杂概念讲清楚。作为浙大 SOFTMAX 宣讲团成员,面向中学生与公共部门人员分享大模型、Agent、AI 治理与数据合规,让技术从术语回到人的理解里。
AI 产品最难的不是模型,是人。团队如何接受 Agent、用户如何建立信任、不同岗位在哪里放弃、对“足够好”的阈值在哪——这些比 benchmark 更值得研究。
合唱和舞台训练让我对协作、节奏和表达有一种身体层面的理解。技术之外,我也珍惜那些关于声音、秩序与共振的经验。
★ 8.0
★ 8.8
★ 8.6
★ 9.1
★ 9.6
★ 8.3
★ 9.3
★ 8.8
★ 9.1
如果你对 AI 产品、模型评测、Agent 工作流、Vibe Coding 或技术如何被更好地表达有相似兴趣,欢迎写信给我。
I reply to every thoughtful message. Don't hesitate to reach out.