刘明Ming Liu

6 年工作经验 · 上海

刘明

亚马逊 AWS 数据科学家 & 算法工程师

企业级 LLM / Agent 系统架构设计与落地交付

关于我

亚马逊 AWS(Amazon Web Services)
2020.07 – 至今
数据科学家 & 算法工程师顾问 · 大模型算法 · 上海

亚马逊 AWS 数据科学家 & 算法工程师 & 全栈开发,专注企业级 LLM / Agent 系统的架构设计与落地交付,拥有完整的 Multi-Agent、RAG 与 Post-training 生产化经验,以及丰富的机器学习、深度学习、强化学习与前后端 / CI-CD 开发经验。

负责企业级 AI / LLM 系统架构设计与应用开发落地,近四年主攻 Agentic RAG、Multi-Agent、复杂任务编排、Text2SQL 与 Post-training;主导亚马逊中国核心大模型平台,建设 To-B MLOps / LLMOps 平台及定制化 multi-agent / post-training 项目,已在金融、制造、医疗、电商、航空、能源等行业规模化落地。

多次担任 Tech Lead / FDE Lead / 总架构师,输出可复用的 Agent 平台能力与后训练方法论,具备从算法研发到商业化交付的端到端负责经验。

核心技术栈

  • Multi-Agent
  • Post-training
  • GraphRAG
  • 强化学习
  • 机器学习 / 深度学习
  • MLOps

服务客户

金融与保险
汇丰银行 · 招商银行 · 国泰世华银行 · 宏利金融(Manulife)
医药与医疗
阿斯利康 · 礼来 · 迈瑞医疗
工业与制造
西门子能源 · 罗伯特·博世 · 宝马集团 · 明阳智慧能源 · 中国钢铁 · 台湾樱花 · 捷普(Jabil)
航空与航运
空中客车 · 国泰航空 · 中国南方航空 · 达飞海运
科技与消费
特斯拉 · OPPO · Shopify · 群晖科技(QNAP)· 安利 · Nothing
公共部门
台湾地区公共部门

项目经历

点击任一项目展开技术细节

  1. Agentic 量化因子研究自动化系统
    2026.04 – 2026.08
    Rogo(美国 · 金融科技)
    算法负责人
    Multi-AgentGraphRAGQuant Research

    为量化因子研究搭建 Multi-Agent 自动化闭环,覆盖数据接入、因子编码、统计评估与自主研究迭代,把因子想法到结论的周期从数天压缩至分钟级。

    • 职责分离的评估机制:LLM 只负责生成因子表达式,所有统计裁决交由确定性评估流程完成,避免模型幻觉影响研究结论。
    • 反数据挖掘体系:将多重检验校正、样本外一次性消费、回测区间固定等规则固化为工程约束并纳入 CI 校验,把过拟合风险从依赖研究员自律转为系统强制。
    • 自主研究闭环:编排因子规划、代码迭代、对抗性评审与探索调度多个 Agent 协作,形成提出 - 回测 - 观测 - 提案的完整循环。
    • 知识库与负知识沉淀:以 GraphRAG 混合召回历史研究结论,将已被证伪的方向作为先验回灌调度器,避免重复挖掘同一路径。
    • 成果:真实数据端到端验证中,系统对无效因子如实给出不放行结论,验证了防假阳能力而非制造好看的回测。
  2. Agentic 自动化投标响应系统与领域模型后训练
    2026.02 – 2026.07
    西门子能源(工业能源)
    技术算法总负责人 · 带领 3 人团队
    Multi-AgentPost-trainingGRPO

    面向工业投标场景交付多 Agent 协作系统与领域模型后训练全流程,覆盖多格式规格书解析到结构化技术参数填写,支持多语言。

    • 多 Agent 协作架构:按规划、文档理解、参数匹配、校验、合规与撰写等职责拆分 Agent,其中匹配与校验环节承担绝大部分模型调用量。
    • 训练数据工程:构建 Agent 轨迹训练集,通过多路径展开、有用性过滤与对抗性负样本提升模型在真实文档上的稳健性。
    • 分阶段后训练:采用课程式训练与强化学习优化,显著提升模型在未见格式上的抽取准确率,并大幅降低幻觉率。
    • 知识边界与拒答:构建多层拒答机制,让模型在证据不足时主动弃答而非编造,显著降低不合理拒答与错误作答。
    • 成果:抽取质量接近前沿通用模型水平,单文档处理成本降低一个数量级,交付周期从人工周级压缩至天级。
  3. Agentic 反欺诈风控系统
    2025.10 – 2026.03
    汇丰银行(金融风控)
    算法负责人 / FDE Lead
    Multi-AgentGNNSequence Modelling

    为信用卡风控设计 Multi-Agent 反欺诈系统,覆盖多源特征编码、集成打分、知识图谱检索与自主调查,将新型欺诈手法从发现到被系统拦截的周期从数周压缩到分钟级。

    • 多源信号融合:将统计特征、交易序列表征与图结构表征统一融合打分,自动捕捉单一维度难以发现的跨模态组合信号;相较原有纯手工特征方案,正负样本分数区分度显著提升,人工复核区间明显收窄。
    • 行为序列建模:以自监督方式学习客户消费行为表征,将长期行为模式压缩为可在线服务的客户向量。
    • 图神经网络团伙识别:在实体级知识图谱上按关系类型分权重建模,使强关联信号不被弱关联稀释,识别多跳范围内的欺诈团伙结构。
    • 自主调查与规则治理:构建分层 Agent 决策架构完成线索调查与规则生成,规则须通过回测门槛并经人工审批方可生效,并设置护栏禁止 Agent 自行臆造风控阈值。
    • 成果:在无标签场景下仅凭行为模式识别出此前未被发现的隐藏欺诈团伙,并在生产规模下带来可观的模型效果提升。
  4. 低延迟实时语音 Agent 系统
    2025.05 – 2025.11
    Dhan(印度 · 金融科技)
    技术与算法总负责人 · 带领 5 人团队
    Realtime VoicePost-trainingTool Calling

    面向金融业务场景交付实时语音 Agent,覆盖语音识别、Agent 决策、工具执行到语音合成的完整链路,支持多语言混说交互。

    • 分层决策架构:以路由层与执行层分离的方式解耦意图识别、任务分发、工具选择与异常恢复,支撑低延迟多轮对话与复杂业务流程。
    • 流式链路延迟优化:对识别、推理与合成三段分别优化并全链路流式化,在高并发压测下将端到端首响延迟控制在秒级。
    • 路由与工具选择后训练:对路由与执行模型分别进行偏好优化与强化学习训练,在保持高路由准确率的同时将不安全路由控制在极低水平,并显著提升复杂场景下的工具选择准确率。
    • 语音噪声鲁棒性:自研识别错误模拟器向训练语料注入噪声与语法扰动,并引入证据校验与恢复轨迹,降低误识别引发的级联失败。
    • 成果:在高并发下稳定保持秒级响应,人工转接率维持在极低水平。
  5. 具长期记忆的自进化 Multi-Agent 系统
    2025.02 – 2026.02
    Shopify(跨境电商)
    产品设计与 Agent 开发总负责人 · 带领 7 人团队
    Long-term MemoryAgent OrchestrationRecommendation

    为内容创作者交付具长期记忆的自进化 Multi-Agent 系统,实现从用户理解、任务规划、Agent 执行到反馈与自我评估的闭环。

    • 动态意图框架:将 Agent 与任务统一映射为多层意图结构,支撑跨 Agent、跨工具的一致路由与调度。
    • 编排与异步协作:支持动态工作流生成、Agent 间异步通信与契约约定、共享工作区、阻塞检测与错误自动恢复,并支持人工介入。
    • 双向表征匹配推荐:将传统推荐需求重构为 Agent 驱动方案,以创作者画像与商品表征双向匹配,并依据实际转化表现周期性协同更新,改善冷启动场景的匹配质量与覆盖度。
    • 可配置 Agent 平台:沉淀策略、提示、工具与验收用例的配置化平台,引导无 AI 背景用户完成能力搭建,实现可复制交付。
    • 事件驱动的主动分析:对接多个电商与社交数据源,围绕互动、转化与增长指标自动分析并触发相应运营动作。
  6. 多意图智能客服 Agent
    2024.01 – 2024.04
    USPACE 悠勢科技(智慧停车)
    Tech Lead
    Multi-intentTool OrchestrationHybrid Retrieval

    为智慧停车平台构建繁体中文多意图客服 Agent,覆盖停车、支付、保险与洗车多条业务线,将大量后端接口收敛为标准化工具集并控制在当期模型的上下文预算内。

    • 多标签意图识别:将单意图分类升级为多标签识别并联合抽取槽位与置信度,解决单轮对话含多诉求时的意图遗漏与参数串联问题。
    • 分级模型路由:以轻量模型完成领域识别与低成本预路由,仅将复杂、写操作或低置信请求升级至高阶模型,大幅降低推理成本。
    • 任务规划与并发执行:在当期模型尚不支持并行工具调用的条件下自研规划与依赖图执行机制,按工具依赖关系自动编排并并发执行无依赖任务。
    • 工具检索与上下文压缩:以业务域筛选叠加混合检索控制单次暴露的工具数量,在保持高召回的前提下将工具上下文压缩至原先的四分之一左右。
    • 写操作安全与澄清体验:为写操作设计幂等、二次确认与补偿机制,可逆操作优先、不可逆操作最后执行;并将缺失参数批量一次询问,明显减少澄清轮次。
    • 成果:基于真实对话构建评测集验证,端到端任务达成率与一次解决率较原有方案大幅提升,人工工单显著下降。
  7. 超大规模时序异常检测平台
    2023.08 – 2023.12
    招商银行(金融运维)
    算法顾问
    Time SeriesAnomaly DetectionAlert Reduction

    为行内数千个应用系统的海量监控序列设计周期基线与深度重构双通道融合的异常检测方案,兼顾实时性与大规模训练的算力可行性。

    • 实时与离线分离架构:流式链路负责实时检测与告警收敛,离线链路负责训练与校准,实时侧延迟控制在秒级。
    • 业务日历感知的周期基线:将银行特有的结算、发薪、节假日与调休因素纳入基线建模,以标准化残差作为异常分数,避免例行业务高峰造成大面积误报。
    • 多指标联合重构:以重构模型联合建模多维指标,识别单指标视角无法发现的指标间关系破裂型异常。
    • 算力可行性工程:对监控指标自动分型,仅对强周期序列使用重量级模型,其余采用轻量方案,并通过序列聚类共享季节性与并行训练,将全量重训压缩到可接受的时间窗内。
    • 全局阈值可分性指标:针对单序列指标表现良好但跨海量指标无法统一设阈的问题,自定义分数可分性度量并以稳健标准化与尾部概率校准优化,大幅减少需人工配置阈值的序列数量。
    • 成果:在多来源评测集上取得良好检测精度与低漏报率;经多点确认、同指标合并与拓扑根因收敛,原始告警量压缩一个数量级以上,无效告警明显下降。
  8. 集装箱智能验箱视觉系统
    2023.07 – 2023.09
    达飞海运 CMA CGM(航运)
    算法负责人
    Instance SegmentationVisual MetrologyIndustry Standards

    构建闸口不停车采集到箱损识别、尺寸测量、标准化编码与修箱估价的完整验箱链路,替代持证验箱员对全量箱的目视初检,人工转为例外复核与抽检。

    • 多视角感知:门架多路相机不停车抓拍,识别箱号并与图像绑定校验,以实例分割输出多类箱体损伤;按每箱误报上限标定工作点,保证可计费损伤的高召回。
    • 单目物理尺寸测量:以行业标准规定的箱体角件间距作为唯一已知尺度,结合相机标定与透视校正实现毫米级损伤尺寸测量;并显式声明测量下限,小于阈值的损伤只作定性标记转人工。
    • 行业标准编码与单据自动化:将识别结果自动映射为航运业标准的组件、损伤、修理码与计量规则,按两套行业阈值给出放行、维修或拒收建议,并直接生成修箱估价单据草稿,大幅减少人工录入。
    • 长尾与跨站点适应:针对夜间、雨天与逆光低质图像做增强建模,对易混淆的锈蚀与污渍类别做难分样本挖掘,新站点以小样本标注加阈值校准上线而非全量重训。
    • 成果:分级结论与持证验箱员判定高度一致,多数空箱实现自动放行,人工复检工作量与估价争议率显著下降。
  9. 多目标柔性车间排程系统
    2023.05 – 2023.07
    捷普 Jabil(电子制造)
    算法工程师
    Combinatorial OptimisationDistributed ComputingMES Integration

    为高混线电子制造车间交付分布式启发式排程系统,覆盖多条产线与全部设备的滚动周排程,支撑插单与设备异常后的班中快速重排。

    • 贴近现场的问题建模:将换型时间与产线工艺相似度关联,并纳入设备能力约束、工序间时窗要求、物料时效限制与换线人力有限带来的双资源约束;硬约束下沉到解码环节强制满足,只把交期与换线成本留在目标函数,避免输出不可执行的排程。
    • 启发式与局部搜索结合:以双段编码配合主动调度解码,叠加针对关键路径的局部搜索,并在种群多样性下降时自适应调整搜索强度。
    • 分布式实现与真实瓶颈定位:采用异构岛模型异步并行以消除同步等待;剖析发现瓶颈在解码而非通信,据此改为共享只读工艺数据、仅传递解方案并向量化解码,显著提升单位时间评估吞吐。
    • 技术选型与基准对照:选择启发式方案的关键理由是随时可中断取解特性与目标函数可随季度调整;同时以精确求解器作为标尺量化解质量差距,并为结果标注下界 gap,而非仅宣称最优。
    • 抑制计划抖动与系统集成:以已下发计划播种初始解并对偏离施加惩罚,避免重排造成现场混乱;对接企业资源与制造执行系统取数回写,并用历史工时反向校准标准工时。
    • 成果:经影子运行与产线对照验证,完工跨度、换线总时长与交期达成率均获明显改善,单次排程耗时压缩至可支撑班中重排的水平。
  10. 私有化部署 AIOps 故障定位 Agent
    2023.02 – 2023.06
    安利中国(消费品)
    算法工程师
    AIOpsLoRA Fine-tuningTool Calling

    为运维团队交付可私有化部署的故障定位 Agent,基于开源模型做领域后训练,编排日志、代码仓库与工单多源取证,自动输出故障时间线、变更嫌疑排序与根因假设。

    • 工具编排与模型分工:利用模型原生工具调用能力编排多个取证工具,无需自研解析层;并按上下文长度与能力差异做双模型分工,规避单一模型的能力限制。
    • 日志证据提取与压缩:以在线日志模板挖掘将海量原始日志收敛为少量模板,再经高召回筛选与模型精排形成可注入上下文的证据链。
    • 按召回优先设计防漏检:筛选阈值以历史根因证据保留率而非精确率调优,每条结论保留可下钻至原始日志的链接,并额外保留统计异常但未纳入证据的清单兜底,避免摘要成为信息黑洞。
    • 变更嫌疑度排序:对故障窗口前的代码变更计算部署时间接近度、调用路径交集、依赖可达性与改动风险面等多维特征加权打分,权重在历史有明确肇事变更的故障上校准;模型只负责解释证据、不改变分数,保证可解释与可回溯。
    • 领域后训练与数据飞轮:以历史复盘文档反向构造训练样本,采用参数高效微调在单卡上完成训练与推理;定位为辅助决策而非自动修复,仅对白名单动作提供一键执行入口并由人确认,故障关闭时自动生成复盘草稿再回流训练语料。
    • 成果:故障定位耗时中位数缩短一半以上,根因假设命中率经双人盲评验证达到可用水平,复盘首稿撰写时间大幅压缩。
  11. 多账户云安全异常检测与告警降噪
    2022.08 – 2022.12
    宝马集团(汽车制造)
    算法负责人
    UEBAWeak SupervisionSecurity Analytics

    为大规模多账户云环境构建安全异常检测与告警降噪系统,按主体类型分群建立行为基线,输出统一风险评分并将孤立告警关联成攻击链。

    • 数据底座与成本控制:将多源安全日志统一为开放标准格式以支撑高效查询,显著降低历史回溯的扫描量与成本;并仅对高风险资源开启细粒度审计,避免全量采集的高额开销。
    • 按主体分群建模:对人类用户与各类自动化角色分别建立行为基线,提取罕见操作、首次访问、跨账户提权与异常访问链路等特征,结合无监督检测、会话序列分析与监督排序输出带证据链与攻击框架标签的告警。
    • 标签稀缺下的严谨评测:结合弱监督、正例无标签学习、历史处置反馈与攻击模拟构造训练数据,以时间切分与主体分组避免数据泄漏;上线前经数周影子运行与原规则系统对照,并刻意避免采用在极端不平衡下容易虚高的评价指标。
    • 闭环处置与数据合规:检测结果回写安全平台并驱动自动化处置;针对当地员工数据治理要求对主体身份做假名化,并采用分层数据保留策略。
    • 成果:在保持已确认安全事件高召回的前提下,进入人工深度调查队列的告警量大幅压降,告警排序质量与平均处置决策效率显著改善。
  12. 家电安装现场质检视觉系统
    2022.07 – 2022.10
    罗伯特·博世(家电制造)
    算法工程师
    Instance SegmentationGraph ReasoningEdge Inference

    面向家电安装完成后的现场质检,建设安装即检的视觉系统,替代覆盖率极低的事后人工抽检。系统按工单与机型下发必拍视角,端侧校验拍摄完整性,云端识别管线与接口状态并返回缺陷说明与整改建议,支持现场复拍闭环。

    • 柔性管线建模:针对管线无固定形态的难点,以实例分割结合骨架化与图结构推理还原管线与接口的连接关系,并利用方向、纹理与直径特征处理交叉、遮挡与断裂情况。
    • 通用模型加规则配置架构:以通用部件模型搭配机型级规则配置支撑数十个机型,新机型通过少量样本微调与阈值校准即可上线,无需重训。
    • 成果:在覆盖数千工单的验证集上取得较高自动判定覆盖率与整体准确率;经按安装人员分组的灰度验证,二次上门返工率相对下降约三分之一,端侧推理延迟满足现场实时校验要求。
  13. 竞品运价变动预测与采集调度优化
    2022.03 – 2022.10
    中国南方航空(民航)
    算法工程师
    Survival AnalysisResource AllocationCausal Care

    将竞品运价采集重构为预算约束下的信息新鲜度优化问题,以动态分配付费查询额度替代固定频率全量轮询,在大幅降低采集成本的同时提升变价发现能力。

    • 生存模型预测变价风险:以离散时间生存模型预测各监控单元下一窗口的变价风险;针对变价发生在两次查询之间、真实时点不可观测的问题采用区间删失建模,而非简单记为下次观测时刻。
    • 优先级索引调度:基于松弛方法构造优先级索引,综合变价概率、预期变动幅度、航线收入权重与当前信息陈旧度分配每日预算;并按收入等级设置最大陈旧度约束,保证长尾航线也被周期性覆盖。
    • 选择偏差治理:预留部分额度做随机探测并采用带随机性的调度,训练与评估引入倾向得分加权,避免只观察模型预期会变价的样本所造成的自我强化偏差;线上以按市场分组的实验设计规避相邻航线与共享预算的干扰。
    • 成果:日均查询量与采集成本显著下降的同时,收入加权的变价召回率保持在高位,信息陈旧度大幅改善,对竞品变价的响应时延缩短至原先的三分之一左右。
  14. 基于强化学习的机票动态定价优化
    2022.03 – 2022.07
    国泰航空(民航)
    算法负责人
    Reinforcement LearningDemand SimulationOffline Evaluation

    在既有收益管理体系之上叠加强化学习定价策略层,不替换原有需求预测、库存控制与网络优化模块;策略仅在既定信任域内输出价格调整量,确保不突破原有收益管理约束。

    • 序贯决策建模与奖励设计:将定价建为有限时域序贯决策问题,状态涵盖库存、时间、预订曲线偏离与竞品价差等;奖励以当期净收入为主体,并引入座位位移的机会成本与起飞时空座惩罚,避免策略过度保守或过度激进。
    • 需求仿真器:以随机到达过程结合离散选择模型联合建模购买、放弃、舱位下沉与竞品替代行为;对舱位关闭导致的需求删失做还原处理,并以滚动起点回测校准不同预订周期的曲线。
    • 内生性处理:以燃油附加费、汇率与促销时点等外生信号识别价格弹性,降低需求高导致价格高所造成的估计偏差,而非直接从历史价量关系拟合弹性。
    • 约束与稳健性:以信任域、正则约束与硬性动作掩码限制策略偏离既有体系,并覆盖价格阶梯、购票规则与监管上限;在弹性置信区间内做随机化训练并以风险度量为稳健目标,降低仿真到现实的迁移风险。
    • 离线评估严谨性:以拟合 Q 评估为主并结合多种离线估计方法交叉验证,对长预订周期下方差过大的估计方法只用于短窗口,避免高估策略收益。
    • 成果:离线仿真与影子模式下单位收益与客座率均获改善,且在悲观弹性假设下仍保持正收益;影子运行显著降低分析师人工干预并大幅提升定价复核覆盖量。
  15. 球场运动员实时多目标追踪系统
    2021.05 – 2021.09
    上海广播电视台(媒体)
    算法工程师
    Multi-Object TrackingInference OptimisationStreaming

    构建足球赛事实时多目标追踪系统,实现球员检测、身份关联、轨迹生成与场上定位,并完成推理性能与流式服务的工程优化以支撑实时转播场景。

    • 追踪模型训练与优化:训练实时多目标追踪模型完成检测与身份关联,并通过多尺度训练、遮挡增强与困难样本挖掘明显降低身份切换错误。
    • 标注效率工程:搭建视频数据清洗与半自动标注系统处理大规模帧与目标实例,将标注周期压缩过半。
    • 推理性能优化:以半精度、推理引擎优化、异步流水线与显存复用将吞吐提升一倍以上、平均延迟减半,满足实时性要求。
    • 流式服务架构:以消息队列解耦视频接入、模型推理与轨迹处理模块,支撑多路高清视频流并行实时分析。
  16. 强化学习自动驾驶平台
    2020.07 – 2021.04
    Amazon DeepRacer(云服务产品)
    算法工程师
    Reinforcement LearningComputer VisionSim-to-Real

    参与单目与双目视觉强化学习算法研发,构建从图像感知到车辆转向与速度控制的端到端驾驶策略,支持云端仿真训练并部署至实体车辆。

    • 端到端策略网络:设计卷积编码器结合演员评论家架构的策略网络,将赛道图像直接映射为连续转向与油门指令,显著提升赛道完成率并降低出界次数。
    • 训练稳定性与收敛加速:针对训练不稳定、探索不足与易陷局部最优等问题,引入优先经验回放、奖励归一化、梯度裁剪与分阶段学习率,并采用最大熵框架提升连续动作空间的探索能力,明显缩短收敛时间并提升样本利用率。
    • 双目视觉融合:完成左右视图的同步预处理与特征融合,利用视差信息增强对赛道边界与前方障碍的空间感知,明显降低动态障碍场景下的碰撞率。
    • 分布式训练与仿真评估:搭建分布式训练与仿真评估流程,支持多环境并行采样、异步训练、检查点管理与自动化评估,显著提升单位时间有效交互样本数并缩短实验周期。
    • 平台服务接口:开发训练任务管理、算法与传感器配置、奖励函数上传、模型评估与导出等对外接口,采用异步任务与幂等控制保证高可用与稳定响应。
    • 仿真到实车迁移:通过图像扰动、纹理随机化与动作平滑约束缩小仿真与真实环境差异,大幅提升实车连续完赛率并降低方向控制抖动。
  17. 强化学习智慧供热系统(共三期)
    2019 – 2026.08
    淄博热力(城市公共事业)
    Tech Lead / 算法负责人
    Reinforcement LearningTime-series ForecastingIndustrial Control

    为城市供热网络构建小时级预测与智能控制系统,覆盖数百座换热站的站级调控与部分小区的房间级供热优化,项目历经三期持续演进。

    • 领域特征体系:构建供热领域特征仓库,持续挖掘气象、建筑热惯性与管网水力工况等特征,为多模型复用提供统一供给。
    • 多模型融合预测:结合树模型与时序预测方法构建小时级站控预测体系,兼顾强周期性负荷基线与突变工况响应。
    • 强化学习控制策略:以强化学习构建供热控制策略,通过奖励设计与调优提升控制稳定性,效果优于人工调控,温度稳定性显著改善。
    • 异常检测与差异化控温:构建实时异常检测识别异常站点并触发差异化控温策略,避免个别站点异常影响全网调节质量。
    • 团队与数据平台:作为 Tech Lead 带领合作团队建设数据平台,支撑缴费、客服与财务等多个业务系统的数字化转型。
    • 成果:年均节能接近一成,居民投诉率大幅下降。

专业技能

LLM / Agent
Multi-Agent Orchestration、Agent Harness、Tool-call / ReAct Agent、动态意图路由、DAG Planner、Agentic RAG、GraphRAG、长期记忆、Guardrails 与可观测(Langfuse / OTel)
Post-training
SFT、LoRA / PEFT、DPO、GRPO、课程学习、拒绝采样、process reward、gradient masking、R-Tuning 拒答、trajectory 数据工程(DFSDT)
强化学习
DDPG、PPO、SAC、CQL、Bandit(UCB)、奖励函数设计、离线策略评估(FQE / MIS / WDR)、CVaR 稳健优化、Sim-to-Real
机器学习 / 深度学习
XGBoost、LightGBM、Prophet、DeepAR、LSTM Autoencoder、Transformer、RGCN / GNN、YOLOv8-seg、FairMOT、生存分析、遗传算法 / Memetic GA
工程与云
Python、PyTorch、TensorFlow、AWS(Bedrock AgentCore、SageMaker、Neptune、Security Lake、RoboMaker)、Ray、Flink、Spark、Kafka、Triton、TensorRT、Docker、PostgreSQL、ROS / Gazebo

本页面为个人介绍,内容不含雇主或客户的保密信息。

6 years of experience · Shanghai

Ming Liu

Data Scientist & Algorithm Engineer, Amazon AWS

Architecting and shipping enterprise LLM / Agent systems

About

Amazon Web Services (AWS)
Jul 2020 – Present
Data Scientist & Algorithm Engineer (Consultant) · Large-model algorithms · Shanghai

Data scientist, algorithm engineer and full-stack developer at Amazon AWS, focused on architecting and delivering enterprise LLM / Agent systems — with production experience across multi-agent orchestration, RAG and post-training, plus a strong background in ML, deep learning, reinforcement learning and full-stack / CI-CD engineering.

I own architecture and delivery of enterprise AI / LLM systems, with the last four years centred on Agentic RAG, multi-agent systems, complex task orchestration, Text2SQL and post-training. I led Amazon China's core large-model platform, building a B2B MLOps / LLMOps platform and bespoke multi-agent / post-training engagements now running at scale in finance, manufacturing, healthcare, e-commerce, aviation and energy.

Repeatedly serving as tech lead, FDE lead and chief architect, I produced reusable agent-platform capability and a post-training playbook, owning delivery end to end from algorithm research through to commercial launch.

Core stack

  • Multi-Agent
  • Post-training
  • GraphRAG
  • Reinforcement Learning
  • ML / Deep Learning
  • MLOps

Clients

Financial services
HSBC · China Merchants Bank · Cathay United Bank · Manulife
Pharmaceuticals & medtech
AstraZeneca · Eli Lilly · Mindray
Industrial & manufacturing
Siemens Energy · Robert Bosch · BMW Group · Mingyang Smart Energy · China Steel · Sakura Taiwan · Jabil
Aviation & shipping
Airbus · Cathay Pacific · China Southern Airlines · CMA CGM
Technology & consumer
Tesla · OPPO · Shopify · QNAP · Amway · Nothing
Public sector
Taiwan public sector

Projects

Select a project to expand the technical detail

  1. Agentic quant factor research automation
    Apr 2026 – Aug 2026
    Rogo (USA · fintech)
    Lead, algorithms
    Multi-AgentGraphRAGQuant Research

    A multi-agent loop automating quant factor research end to end — data ingestion, factor encoding, statistical evaluation and autonomous iteration — cutting the cycle from idea to conclusion from days to minutes.

    • Separation of duties. The LLM only authors factor expressions; every statistical verdict is made by a deterministic evaluation path, so model hallucination cannot reach the conclusion.
    • Anti-data-mining controls. Multiple-testing correction, single-consumption hold-out and fixed backtest windows are encoded as engineering constraints and enforced in CI, shifting overfitting risk from researcher discipline to something the system structurally prevents.
    • Autonomous research loop. Factor planning, code iteration, adversarial review and exploration scheduling are orchestrated across agents into a propose-backtest-observe-submit cycle.
    • Knowledge base with negative results. GraphRAG hybrid retrieval surfaces prior findings, feeding already-falsified directions back into the scheduler as priors so the same path is not re-mined.
    • Outcome. In end-to-end validation on real data the system correctly refused a factor without genuine alpha — demonstrating false-positive resistance rather than manufacturing a flattering backtest.
  2. Agentic bid-response system with domain post-training
    Feb 2026 – Jul 2026
    Siemens Energy (industrial energy)
    Overall technical lead · team of 3
    Multi-AgentPost-trainingGRPO

    Delivered a multi-agent system and the full domain post-training pipeline for industrial bid response, from multi-format specification parsing through to structured technical parameter completion, across multiple languages.

    • Multi-agent architecture. Responsibilities split across planning, document understanding, parameter matching, verification, compliance and drafting, with matching and verification carrying the bulk of model calls.
    • Training data engineering. Built an agent-trajectory training set using multi-path expansion, usefulness filtering and adversarial negatives to harden the model against real-world documents.
    • Staged post-training. Curriculum training combined with reinforcement learning materially improved extraction accuracy on unseen formats and sharply reduced hallucination.
    • Knowledge boundaries and refusal. A layered refusal mechanism makes the model decline when evidence is insufficient rather than invent an answer, cutting both unjustified refusals and wrong answers.
    • Outcome. Extraction quality approaches that of frontier general models, cost per document fell by an order of magnitude, and turnaround compressed from person-weeks to days.
  3. Agentic fraud detection system
    Oct 2025 – Mar 2026
    HSBC (financial risk)
    Lead, algorithms / FDE lead
    Multi-AgentGNNSequence Modelling

    A multi-agent fraud detection system for credit-card risk spanning multi-source feature encoding, ensemble scoring, knowledge-graph retrieval and autonomous investigation — shortening the lag between a novel fraud tactic being identified and being blocked from weeks to minutes.

    • Multi-source signal fusion. Statistical features, transaction-sequence representations and graph-structure representations are fused into a single score, capturing cross-modal combinations no single view reveals. Against the incumbent hand-crafted-feature approach, separation between fraudulent and legitimate scores improved substantially and the manual-review band narrowed.
    • Behavioural sequence modelling. Self-supervised training learns representations of customer spending behaviour, compressing long-run patterns into a customer vector servable online.
    • Graph neural network ring detection. Relation-typed weighting over an entity knowledge graph keeps strong associations from being diluted by weak ones, exposing fraud-ring structure several hops out.
    • Autonomous investigation with rule governance. A layered agent architecture handles lead investigation and rule generation; a rule must clear a backtest threshold and human approval before taking effect, and guardrails forbid agents from inventing risk thresholds.
    • Outcome. With no labels available, behavioural patterns alone surfaced previously undetected fraud rings, with a meaningful model-performance lift projected at production scale.
  4. Low-latency real-time voice agent
    May 2025 – Nov 2025
    Dhan (India · fintech)
    Overall technical lead · team of 5
    Realtime VoicePost-trainingTool Calling

    A real-time voice agent for financial services covering speech recognition, agent decision-making, tool execution and speech synthesis, handling code-mixed multilingual conversation.

    • Layered decision architecture. Separating routing from execution decouples intent recognition, dispatch, tool selection and error recovery, supporting low-latency multi-turn dialogue over complex workflows.
    • Streaming latency optimisation. Recognition, inference and synthesis were each optimised and the whole chain made streaming, holding end-to-end time-to-first-response within a second or so under high-concurrency load.
    • Post-training for routing and tool selection. Preference optimisation and reinforcement learning were applied separately to the routing and execution models, keeping routing accuracy high while holding unsafe routing to a very low rate and markedly improving tool selection on complex cases.
    • Robustness to speech noise. A purpose-built recognition-error simulator injects noise and grammatical perturbation into training data, while evidence checking and recovery trajectories reduce cascading failures from misrecognition.
    • Outcome. Sub-second-scale response held steady under high concurrency, with human handoff kept to a very low rate.
  5. Self-evolving multi-agent system with long-term memory
    Feb 2025 – Feb 2026
    Shopify (cross-border e-commerce)
    Overall lead, product design and agent development · team of 7
    Long-term MemoryAgent OrchestrationRecommendation

    A self-evolving multi-agent system with long-term memory for content creators, closing the loop from understanding the user through planning, execution, feedback and self-evaluation.

    • Dynamic intent framework. Agents and tasks map into one multi-layer intent structure, giving consistent routing and scheduling across agents and tools.
    • Orchestration and async collaboration. Dynamic workflow generation, asynchronous inter-agent communication and contracts, a shared workspace, blocker detection and automatic error recovery, with human intervention supported throughout.
    • Two-sided representation matching. Reframed a conventional recommender brief as an agent-driven design matching creator profiles against product representations, refreshing both on a cadence driven by realised conversion to improve cold-start match quality and coverage.
    • Configurable agent platform. A platform for policies, prompts, tools and acceptance cases that guides users with no AI background through building capability, making delivery repeatable.
    • Event-driven proactive analysis. Integrates multiple commerce and social data sources, analysing engagement, conversion and growth metrics automatically and triggering the corresponding operational actions.
  6. Multi-intent customer service agent
    Jan 2024 – Apr 2024
    USPACE (smart parking)
    Tech lead
    Multi-intentTool OrchestrationHybrid Retrieval

    A traditional-Chinese multi-intent service agent for a smart-parking platform spanning parking, payments, insurance and car washing, consolidating a large backend API surface into a standardised tool set within the context budget of the models available at the time.

    • Multi-label intent recognition. Upgrading from single-intent classification to multi-label recognition with joint slot and confidence extraction fixed missed intents and parameter bleed when one turn carried several requests.
    • Tiered model routing. A lightweight model handles domain detection and cheap pre-routing, escalating only complex, write or low-confidence requests to the premium model — cutting inference cost substantially.
    • Task planning and concurrent execution. Because the models of that era could not call tools in parallel, I built planning and dependency-graph execution that derives ordering from tool dependencies and runs independent tasks concurrently.
    • Tool retrieval and context compression. Domain filtering plus hybrid retrieval limits how many tools are exposed per call, compressing tool context to roughly a quarter of its original size while keeping recall high.
    • Write safety and clarification experience. Write operations carry idempotency, explicit confirmation and compensation, with reversible steps first and irreversible ones last; missing parameters are batched into a single question, cutting clarification turns noticeably.
    • Outcome. Validated on an evaluation set built from real conversations, end-to-end task completion and first-contact resolution rose sharply against the incumbent solution and manual tickets fell significantly.
  7. Large-scale time-series anomaly detection platform
    Aug 2023 – Dec 2023
    China Merchants Bank (IT operations)
    Algorithms consultant
    Time SeriesAnomaly DetectionAlert Reduction

    A dual-channel anomaly detection scheme — seasonal baselines fused with deep reconstruction — for the vast monitoring series behind thousands of banking applications, balancing real-time response against the compute realities of training at that scale.

    • Separated real-time and offline paths. The streaming path handles live detection and alert consolidation while the offline path trains and calibrates, holding real-time latency to seconds.
    • Business-calendar-aware baselines. Settlement cycles, paydays, holidays and substituted working days specific to banking are modelled into the baseline, scoring anomalies from standardised residuals so routine peaks stop generating mass false alarms.
    • Joint multi-metric reconstruction. A reconstruction model over several metrics together catches broken inter-metric relationships that no single-metric view reveals.
    • Making the compute feasible. Automatic series typing sends only strongly seasonal series to the heavyweight model and lighter methods handle the rest; shared seasonality across clustered series plus parallel training compressed full retraining into an acceptable window.
    • A global separability metric. Addressing the gap where per-series metrics look strong yet no single threshold works across a huge metric population, I defined a score-separability measure and optimised it via robust standardisation and tail-probability calibration — sharply reducing the number of series needing hand-set thresholds.
    • Outcome. Good detection accuracy and a low miss rate on a multi-source evaluation set; consecutive-point confirmation, same-metric merging and topology-based root-cause consolidation compressed raw alert volume by more than an order of magnitude, with useless alerts down markedly.
  8. Container inspection vision system
    Jul 2023 – Sep 2023
    CMA CGM (shipping)
    Lead, algorithms
    Instance SegmentationVisual MetrologyIndustry Standards

    An inspection pipeline from drive-through capture at the gate to damage recognition, dimensional measurement, standardised coding and repair estimation — replacing certified surveyors' visual first pass on every container, with people moving to exception review and sampling.

    • Multi-view perception. Gantry cameras capture without stopping the truck, the container number is read and validated against the images, and instance segmentation outputs several damage classes. Calibrating the operating point to a per-container false-positive ceiling secures high recall on billable damage.
    • Monocular physical metrology. Using the industry-standard corner-casting spacing as the only known scale, camera calibration and perspective correction yield millimetre-level damage measurement. A measurement floor is stated explicitly: damage below it is flagged qualitatively for a human rather than given a number.
    • Standards-based coding and document automation. Detections map automatically onto the shipping industry's component, damage and repair codes with their measurement rules, two industry threshold sets yield a release, repair or reject recommendation, and a repair-estimate document is drafted directly — greatly reducing manual entry.
    • Long tail and cross-site adaptation. Augmentation handles night, rain and backlit captures, hard-negative mining separates easily confused corrosion and staining classes, and a new site goes live on few-shot labelling plus threshold calibration rather than full retraining.
    • Outcome. Grading agrees closely with certified surveyors, most empty containers clear automatically, and both manual re-inspection workload and estimate disputes fell significantly.
  9. Multi-objective flexible shop scheduling system
    May 2023 – Jul 2023
    Jabil (electronics manufacturing)
    Algorithm engineer
    Combinatorial OptimisationDistributed ComputingMES Integration

    A distributed heuristic scheduler for a high-mix electronics plant, producing rolling weekly schedules across multiple lines and all equipment, and fast in-shift rescheduling after rush orders or equipment failures.

    • Modelling that matches the floor. Changeover time is tied to process similarity between adjacent orders, alongside equipment capability limits, inter-operation time windows, material shelf-life constraints and dual-resource constraints from limited changeover labour. Hard constraints are pushed into decoding and enforced there, leaving only due dates and changeover cost in the objective — so the scheduler cannot emit a plan the floor can't run.
    • Heuristics with local search. Two-segment encoding feeds an active-schedule decoder, augmented by local search along the critical path, with search intensity adapting when population diversity falls.
    • Distribution, and finding the real bottleneck. A heterogeneous island model runs asynchronously to eliminate sync waits. Profiling located the bottleneck in decoding rather than communication, so process data became read-only shared state, only solutions crossed the wire, and decoding was vectorised — sharply raising evaluation throughput per unit time.
    • Technology choice, measured against a baseline. The heuristic was chosen for its anytime property and for objectives that shift quarterly. An exact solver served as a yardstick to quantify the quality gap, and results carry a lower-bound gap annotation rather than a bare claim of optimality.
    • Damping schedule nervousness, and integration. The previously released plan seeds the initial solution and deviation from it is penalised, so rescheduling does not disrupt the floor. It reads from and writes back to the enterprise resource and manufacturing execution systems, and reverse-calibrates standard times from historical durations.
    • Outcome. Validated through shadow running and line-level comparison, makespan, total changeover time and on-time delivery all improved clearly, and a scheduling run became fast enough to support in-shift rescheduling.
  10. On-premise AIOps incident-triage agent
    Feb 2023 – Jun 2023
    Amway China (consumer goods)
    Algorithm engineer
    AIOpsLoRA Fine-tuningTool Calling

    An on-premise incident-triage agent for the SRE team: an open-source model post-trained on domain data orchestrates evidence gathering across logs, code repositories and tickets, then emits an incident timeline, ranked suspect changes and root-cause hypotheses.

    • Tool orchestration and model division of labour. The model's native tool-calling drives several evidence tools with no bespoke parser, and two model variants split the work by context length and capability to work around the limits of either alone.
    • Log evidence extraction and compression. Online template mining collapses vast raw logs into a small template set, then high-recall filtering and model reranking produce an evidence chain small enough to fit in context.
    • Recall-first, to guard against misses. Filter thresholds are tuned on whether historical root-cause evidence survives rather than on precision; every conclusion links back to the original logs, and statistically anomalous items excluded from the evidence are retained as a backstop so summarisation never becomes an information black hole.
    • Ranking suspect changes. Code changes before the incident are scored on deploy-time proximity, call-path intersection, dependency reachability and blast radius, with weights calibrated on historical incidents having a confirmed culprit change. The model only explains the evidence and never alters the score, keeping the ranking explainable and auditable.
    • Domain post-training and a data flywheel. Training samples were constructed backwards from historical postmortems, with parameter-efficient fine-tuning keeping training and inference on a single GPU. It is positioned as decision support rather than auto-remediation: only allowlisted actions get a one-click entry point and a human confirms, while closing an incident auto-drafts the postmortem which then re-enters the training pool.
    • Outcome. Median time to identify an incident more than halved, root-cause hypothesis accuracy reached a usable level under double-blind review, and drafting a postmortem became far quicker.
  11. Multi-account cloud security detection and alert reduction
    Aug 2022 – Dec 2022
    BMW Group (automotive)
    Lead, algorithms
    UEBAWeak SupervisionSecurity Analytics

    Security anomaly detection and alert reduction across a large multi-account cloud estate: behavioural baselines per principal type, a unified risk score, and isolated findings stitched into attack chains.

    • Data foundation and cost control. Multi-source security logs were normalised to an open standard format for efficient querying, cutting scan volume and cost for historical backfill substantially, while fine-grained auditing was enabled only on high-risk resources to avoid the expense of capturing everything.
    • Cohort-specific modelling. Separate baselines for human users and the various automation roles, with features covering rare operations, first-time access, cross-account privilege escalation and unusual access paths. Unsupervised detection, session-sequence analysis and supervised ranking combine to emit alerts carrying an evidence chain and attack-framework tags.
    • Rigorous evaluation despite scarce labels. Training data came from weak supervision, positive-unlabelled learning, historical dispositions and attack simulation, with temporal splits and principal grouping preventing leakage. Weeks of shadow running compared it against the incumbent rule system, and I deliberately avoided the evaluation metrics that flatter models under extreme class imbalance.
    • Closed-loop response and data compliance. Detections write back to the security platform and drive automated response; to meet local employee-data governance, principal identities are pseudonymised under a tiered retention policy.
    • Outcome. Holding high recall on confirmed incidents, alerts reaching the deep-investigation queue fell substantially, while alert ranking quality and average triage efficiency both improved markedly.
  12. On-site installation QC vision system
    Jul 2022 – Oct 2022
    Robert Bosch (home appliances)
    Algorithm engineer
    Instance SegmentationGraph ReasoningEdge Inference

    A check-on-installation vision system for home-appliance fitting, replacing after-the-fact manual sampling that covered very little. It issues required camera angles per work order and model, validates capture completeness on-device, inspects hoses and connections in the cloud, and returns defect explanations with remediation guidance so the technician can re-shoot on the spot.

    • Modelling flexible hoses. Because hoses have no fixed shape, instance segmentation combined with skeletonisation and graph reasoning recovers which hose connects to which port, with orientation, texture and diameter features resolving crossings, occlusion and apparent breaks.
    • Shared model with per-model rules. A shared component model plus model-level rule configuration supports dozens of variants; a new model goes live on few-shot fine-tuning and threshold calibration without retraining.
    • Outcome. On a validation set spanning thousands of work orders it achieved high automatic-adjudication coverage and overall accuracy; a staged rollout grouped by technician cut repeat-visit rework by roughly a third in relative terms, with on-device inference fast enough for real-time capture validation.
  13. Competitor fare-change prediction and polling optimisation
    Mar 2022 – Oct 2022
    China Southern Airlines (aviation)
    Algorithm engineer
    Survival AnalysisResource AllocationCausal Care

    Reframed competitor fare collection as budget-constrained information-freshness optimisation, replacing fixed-frequency full polling with dynamically allocated paid-query budget — cutting collection cost sharply while improving change detection.

    • Survival modelling of change risk. A discrete-time survival model predicts each unit's hazard of a fare change in the next window. Because changes occur between polls and the true moment is unobservable, the model uses interval censoring rather than naively recording the change at the next observation.
    • Priority-index scheduling. A relaxation-based priority index allocates the daily budget across change probability, expected magnitude, route revenue weight and current staleness, while maximum-staleness constraints by revenue tier guarantee long-tail routes are still covered periodically.
    • Managing selection bias. A share of budget is reserved for random probing with deliberate randomness in scheduling, and propensity weighting enters both training and evaluation — countering the self-reinforcing bias of only observing units the model already expects to change. Online tests were grouped by market to avoid interference between adjacent routes sharing a budget.
    • Outcome. Daily query volume and collection cost both fell significantly while revenue-weighted recall of fare changes stayed high, staleness improved substantially, and response latency to competitor moves dropped to roughly a third of its former level.
  14. Reinforcement-learning dynamic airfare pricing
    Mar 2022 – Jul 2022
    Cathay Pacific (aviation)
    Lead, algorithms
    Reinforcement LearningDemand SimulationOffline Evaluation

    A reinforcement-learning pricing layer added on top of the existing revenue-management stack without replacing demand forecasting, inventory control or network optimisation. The policy only emits adjustments within a defined trust region, so it cannot breach established revenue-management constraints.

    • Sequential decision modelling and reward design. Pricing is framed as a finite-horizon sequential decision problem with state covering inventory, time, booking-curve deviation and competitor gaps. Reward centres on period net revenue, with seat-displacement opportunity cost and a penalty for seats left empty at departure keeping the policy neither too conservative nor too aggressive.
    • Demand simulator. A stochastic arrival process with a discrete-choice model jointly captures purchase, no-purchase, cabin buy-down and competitor substitution. Demand censored by closed cabins is reconstructed, and rolling-origin backtests calibrate curves across booking horizons.
    • Handling endogeneity. Exogenous signals such as fuel surcharges, exchange rates and promotion timing identify price elasticity, mitigating the bias from high demand causing high prices — rather than fitting elasticity straight from the historical price-volume relationship.
    • Constraints and robustness. A trust region, regularisation and hard action masks limit divergence from the incumbent system while respecting fare ladders, ticketing rules and regulatory caps. Randomised training across the elasticity confidence interval with a risk-sensitive objective reduces sim-to-real transfer risk.
    • Rigorous offline evaluation. Fitted Q evaluation is the primary estimator, cross-validated against several alternatives; estimators whose variance explodes over long booking horizons are confined to short windows so policy value is not overstated.
    • Outcome. Unit revenue and load factor both improved in offline simulation and shadow mode, remaining positive even under pessimistic elasticity assumptions; shadow running markedly reduced analyst intervention while greatly increasing pricing-review coverage.
  15. Real-time multi-object player tracking system
    May 2021 – Sep 2021
    Shanghai Media Group (broadcast media)
    Algorithm engineer
    Multi-Object TrackingInference OptimisationStreaming

    A real-time multi-object tracking system for football broadcasts — player detection, identity association, trajectory generation and pitch localisation — with the inference and streaming engineering needed to run live.

    • Tracker training and tuning. Trained a real-time multi-object tracker for detection and identity association, with multi-scale training, occlusion augmentation and hard-example mining noticeably reducing identity switches.
    • Annotation efficiency. Built a video cleaning and semi-automatic annotation system handling large volumes of frames and object instances, more than halving annotation time.
    • Inference optimisation. Half precision, inference-engine optimisation, an asynchronous pipeline and memory reuse more than doubled throughput and halved mean latency, meeting the real-time requirement.
    • Streaming architecture. A message queue decouples ingest, inference and trajectory processing, supporting several concurrent high-definition streams in real time.
  16. Reinforcement learning autonomous driving platform
    Jul 2020 – Apr 2021
    Amazon DeepRacer (cloud product)
    Algorithm engineer
    Reinforcement LearningComputer VisionSim-to-Real

    Worked on monocular and stereo vision reinforcement learning, building end-to-end driving policies from image perception through to steering and speed control, trained in cloud simulation and deployed to physical vehicles.

    • End-to-end policy network. A convolutional encoder with an actor-critic architecture maps track imagery directly to continuous steering and throttle commands, markedly raising lap completion and reducing off-track excursions.
    • Training stability and faster convergence. Against instability, weak exploration and local optima: prioritised replay, reward normalisation, gradient clipping and staged learning rates, plus a maximum-entropy framework to improve exploration of the continuous action space — noticeably shortening convergence and improving sample efficiency.
    • Stereo vision fusion. Synchronised preprocessing and feature fusion across left and right views use disparity to sharpen spatial awareness of track edges and obstacles ahead, clearly reducing collisions in dynamic-obstacle scenarios.
    • Distributed training and simulation evaluation. Built a distributed training and evaluation pipeline with parallel sampling across environments, asynchronous training, checkpoint management and automated assessment — substantially increasing useful interaction samples per unit time and shortening experiment cycles.
    • Platform APIs. Built external interfaces for training-job management, algorithm and sensor configuration, reward-function upload, evaluation and model export, using asynchronous tasks and idempotency control for high availability and stable response.
    • Sim-to-real transfer. Image perturbation, texture randomisation and action-smoothness constraints narrowed the reality gap, substantially raising consecutive real-vehicle completion and reducing steering jitter.
  17. Reinforcement-learning smart district heating (three phases)
    2019 – Aug 2026
    Zibo District Heating (municipal utility)
    Tech lead / lead, algorithms
    Reinforcement LearningTime-series ForecastingIndustrial Control

    Hourly forecasting and intelligent control for a municipal heating network, covering station-level regulation across hundreds of heat exchange stations and room-level optimisation in selected communities, evolving over three phases.

    • Domain feature system. A heating-domain feature store continuously mines weather, building thermal inertia and network hydraulic conditions, giving multiple models one consistent supply.
    • Ensemble forecasting. Tree-based models combined with time-series methods produce hourly station-control forecasts, balancing the strongly seasonal load baseline against response to abrupt operating changes.
    • Reinforcement-learning control. A reinforcement-learning control policy, with reward design and tuning for stability, outperformed manual regulation and improved temperature stability substantially.
    • Anomaly detection and differential control. Real-time detection identifies misbehaving stations and triggers differential control, so an individual station's fault cannot degrade network-wide regulation.
    • Team and data platform. As tech lead, drove a partner team to build the data platform underpinning digital transformation across billing, customer service and finance.
    • Outcome. Annual energy savings approaching ten per cent, with resident complaints down sharply.

Skills

LLM / Agent
Multi-agent orchestration, agent harness, tool-call / ReAct agents, dynamic intent routing, DAG planners, Agentic RAG, GraphRAG, long-term memory, guardrails and observability (Langfuse / OTel)
Post-training
SFT, LoRA / PEFT, DPO, GRPO, curriculum learning, rejection sampling, process reward, gradient masking, R-Tuning refusal, trajectory data engineering (DFSDT)
Reinforcement learning
DDPG, PPO, SAC, CQL, bandits (UCB), reward design, offline policy evaluation (FQE / MIS / WDR), CVaR robust optimisation, sim-to-real
ML / Deep learning
XGBoost, LightGBM, Prophet, DeepAR, LSTM autoencoders, Transformers, RGCN / GNN, YOLOv8-seg, FairMOT, survival analysis, genetic / memetic algorithms
Engineering & cloud
Python, PyTorch, TensorFlow, AWS (Bedrock AgentCore, SageMaker, Neptune, Security Lake, RoboMaker), Ray, Flink, Spark, Kafka, Triton, TensorRT, Docker, PostgreSQL, ROS / Gazebo