Agent-Native platform with claim-based task dispatching, pre-qualification gating, and heterogeneous role collaboration across annotator and QA agents.
抢注式任务分发 + 资格门控 + 异构角色协作,Plan-Act + Reflection 推理范式,Workflow + Agent 混合架构。
AI Product Manager
Agent & Data Strategy
Building AI-Native products at Baidu — Agent platforms, multi-scenario evaluation systems, and training-data strategy across 15+ business lines.
Currently leading Agent platforms, multi-scenario evaluation systems, and training-data strategy at Baidu — covering medical agents, virtual-companion agents, and 15+ business lines.
I practice an AI Native methodology: Agent as a first-class citizen, Specification-Driven, Workflow + Agent hybrid, Evaluation-First, Vibe Coding.
A compact case index: each row exposes the outcome first, then expands into the system logic behind it.
Agent-Native platform with claim-based task dispatching, pre-qualification gating, and heterogeneous role collaboration across annotator and QA agents.
抢注式任务分发 + 资格门控 + 异构角色协作,Plan-Act + Reflection 推理范式,Workflow + Agent 混合架构。
Five-stage evaluation pipeline: wiki, generation, refine, eval, analyse. Wiki-RAG enhanced Rubric generation plus Propose-Evaluate-Revise self-iteration.
医疗问诊 Agent 五阶段解耦评测 Pipeline,覆盖 50+ 临床病种,Cohen’s κ = 0.78。
Specification-Driven evaluation for virtual companion dialogue agents across WenXiaoYan, Shoubai, and in-car products.
-1/0/1 Likert scoring + 11-dimension Analytic Rubric; 21+ batch runs; label consistency 97%。
End-to-end automated pipeline for SFT / DPO / preference data across 15+ business lines, with Bad Case feedback loops as a standard correction mechanism.
线上 Bad Case 回流 → 归因 → 数据补充,面向概率性 Agent 产品建立标准纠错机制。
Full-stack HR operations platform for data-annotation teams: nine-dimension roster management, cross-project dispatch with approval flows and timelines, unified person-day performance metrics, and natural-language AI queries. React 18 + Express 5 + Prisma 6 + PostgreSQL 16, shipped via Docker Compose.
面向数据标注团队的项目人力看板与协同平台——花名册、跨项目调度审批、绩效与成长值统计、AI 自然语言查询、三级 RBAC 权限体系。React + Express + Prisma + PostgreSQL,Docker Compose 部署。
Three-layer evaluation for a micro-expression labeling system (spec v2.7): S0–S5 clip screening with hard filters and VLM quality gates, IAA-based spec stability, VLM-vs-human pipeline accuracy, and a label→generate→relabel semantic-fidelity loop, backed by a seven-factor ablation framework.
微表情标签体系三层评测——S0–S5 视频片段筛选(硬筛 + VLM 质量门)、规范稳定性(IAA)、pipeline 准确度(VLM vs 人工)与端到端语义保真闭环,配套 7 因子消融实验框架。
Professional tarot-reading MVP — a GSAP-driven shuffle-and-draw ritual, a rules engine computing elemental dignities and numeric arcs for free readings, two-tier paid AI readings via DeepSeek, and a complete international payment loop with Creem, Cloudflare Functions and KV.
专业塔罗阅读 MVP——GSAP 洗牌选牌仪式,基于元素尊卑与数字弧的规则引擎驱动免费解读,DeepSeek 双档位付费 AI 解读,Creem + Cloudflare Functions + KV 跑通完整海外收款链路。
A multi-agent system modeling real companies — AI agents communicate peer-to-peer through Feishu/Lark with LLM-planned discussion phases.
A knowledge base for LLM evaluation: methodologies, benchmarks, model comparisons, industry trends, and hands-on experience.
面向大模型评测工作的知识库:评测方法论、基准测试、模型对比、行业动态和评测实践经验。
Daily AI intelligence digest — automatically fetches trending AI projects from GitHub, performs deep analysis via LLMs, and delivers daily digest emails.
每日 AI 情报摘要系统:自动获取 AI 领域热点项目,通过大模型深度分析后发送每日情报邮件。
The original card grid is reframed as an operating model: specs define good work, agents execute open-ended work, workflows keep deterministic control, and evaluation closes the loop.