先读这个
🤖 Agent 开发
造「能自己规划、调用工具、完成任务」的系统。
- 工作对象:Agent 运行时、编排、记忆、工具/MCP、RAG
- 核心难点:可靠性、成本、延迟、上线后的可观测
- 关键词:LangGraph、function calling、多 Agent、ReAct
🧪 AI 评测
回答「它到底好不好、能不能上线」。
- 工作对象:评测集、rubric、LLM-as-judge、回归门禁、红队、RL 环境
- 核心难点:怎么把「好」量化,且不被指标骗
- 关键词:golden set、step-level eval、reward hacking
两个方向正在合流:Agent 岗位里有 9/13 把评测、基准或回归写成明确要求(严格口径,见「高频能力」)。所以评测不是另一个工种,而是 Agent 工程的一部分。
怎么读每张卡
- 点击卡片展开:职责、要求、加分项都已翻译成中文,专有名词保留英文。
- 分析框里的「你已经有的」只引用你
profile.md里写过的经历,没有编造。 - 页面最底部的卡片里可以折叠查看英文原文,方便你练英语时对照。
匹配度含义
岗位速览
-
★★★★★Agentic AI Engineer (LLM)VinFast · 越南 · 河内Agent 开发需到岗海外
-
★★★★★AI Engineer — Agentic AI PlatformAwear.ai · 美国及部分美洲国家Agent 开发限定国家💰 $120K–$200K
-
★★★★★AI Engineer — Hermes Agentpst.ag · 纯远程Agent 开发限定国家
-
★★★★★Senior AI Engineer(多 Agent IT 运维平台)Mapgenesys · 美国芝加哥Agent 开发限定国家💰 $80K–$150K
-
★★★★★AI Agent & Automation DeveloperCherry Bekaert · 美国Agent 开发限定国家💰 $117,400–$172,500
-
★★★★★Senior AI Engineer(临床 AI Agent)Cadence · 美国Agent 开发限定国家💰 基础薪资 $180,000–$245,000
-
★★★★★Principal AI Agent Development Engineer / ArchitectBybit · 马来西亚 · 吉隆坡Agent 开发需到岗海外
-
★★★★★Principal AI Engineer, AI Agent DevelopmentOKX · 香港 / 新加坡Agent 开发需到岗海外
-
★★★★★Principal Engineer, Agent Infrastructure & Memory ArchitectureOKX · 香港 / 新加坡Agent 开发需到岗海外
-
★★★★★Python Backend Development Talent with RAG and Agentic AI ExperienceToptal(客户项目) · 仅限亚洲Agent 开发地区开放
-
★★★★★Backend Engineer, AI (Agent Systems)ActAI · 新加坡Agent 开发地点待确认
-
★★★★★Senior AI Engineer(前线交付方向)H2O.ai · 新加坡Agent 开发需到岗海外
-
★★★★★AI EngineerDistyl AI · 英国 · 伦敦Agent 开发需到岗海外
-
★★★★★GenAI / Agentic AI Evaluation Engineer(质量、安全与可靠性)EY(安永)GDS Assurance · 印度 · 诺伊达AI 评测需到岗海外
-
★★★★★Senior LLM Evaluation Engineer(AI Quality Engineer)Aspire(IT 服务,约旦公司) · 埃及AI 评测需到岗海外
-
★★★★★AI Evaluation EngineerDialpad · 加拿大 · 安大略省基奇纳AI 评测需到岗海外💰 安大略省基础薪资 CAD $96,000–$116,250
-
★★★★★Member of Technical Staff, North Modelling (Evals)Cohere · 英国 · 伦敦AI 评测需到岗海外
-
★★★★★Senior Research Scientist, Model EvaluationCohere · 英国 · 伦敦AI 评测需到岗海外
-
★★★★★AI Quality EngineerDealstitch AI · 美国AI 评测限定国家
-
★★★★★Open Source AI Engineer (TypeScript)Arize AI · 全球远程AI 评测地区开放💰 $185,000–$200,000
-
★★★★★Research Engineer, Coding Evaluation & Training DataSurge AI · 地点未写AI 评测地点待确认
-
★★★★★Backend Engineer(RL 环境 / 训练基础设施)Surge AI · 地点未写AI 评测地点待确认
-
★★★★★RL Environments ArchitectSurge AI · 地点未写AI 评测地点待确认
-
★★★★★Full Stack Engineer(评测与标注平台)Surge AI · 地点未写AI 评测地点待确认
-
★★★★★Research Engineer, Machine Learning (Reinforcement Learning)Anthropic · 英国 · 伦敦AI 评测需到岗海外💰 £260,000–£630,000
-
★★★★★AI Red Team Lead EngineerU.S. Bank(美国银行) · 美国 · 夏洛特AI 评测需到岗海外💰 $126,000–$149,000
-
★★★★★Machine Learning Engineer, Global Public SectorScale AI · 英国 · 伦敦AI 评测限定国家💰 $80,000–$120,000
🤖 Agent 开发(13)
从「一个人扛起整个 Agent 层」的创业公司,到大厂的 Agent 平台和前线交付岗位。
Agent 开发需到岗海外★★★★★
Agentic AI Engineer (LLM)
VinFast · 越南 · 河内
Agent 编排RAGfunction callingLoRA/DPO 微调vLLM
主要职责
- 设计、开发、优化基于 LLM 的 AI Agent,面向真实场景,重点看准确率、延迟、可扩展性和可靠性
- 调研、评估并集成最先进的基础模型(OpenAI、Gemini、Claude、开源 LLM 等)
- 用 SFT、LoRA/QLoRA、DPO、偏好优化、模型蒸馏等技术对大模型做微调、对齐和优化
- 设计多 Agent 系统、Agent 编排、规划、记忆、工具调用和工作流执行架构
- 构建 RAG 流水线:检索、重排(rerank)、索引、Embedding 优化、知识集成
- 设计并优化 function calling 流水线,用于结构化任务执行和外部工具集成
- 为 AI 系统建立评测框架:自动化基准测试、离线评测、线上 A/B、幻觉检测
- 通过 prompt 工程、模型路由、缓存、批处理、量化、投机解码和 serving 优化来提升推理性能
- 开发并维护 API 与可复用的 AI 服务,方便产品快速集成
- 为生产环境里的 Agent 建设可观测能力:日志、链路追踪、监控、性能分析
- 持续跟踪新模型和新 Agent 框架,快速验证并落地到生产
- 与产品、后端、前端、基础设施团队协作,交付生产级 AI 功能;产出技术文档、架构设计和可复用组件
- 指导初级工程师,做技术评审,参与制定工程规范和 AI 最佳实践
任职要求
- 计算机、AI、机器学习、数据科学、软件工程等相关专业本科或硕士
- 3 年以上 AI/ML 工程、NLP、LLM 应用或相关软件工程经验(按职级调整)
- 深入理解 LLM、Transformer 架构和现代生成式 AI 技术
- 用 LangGraph、LangChain、LlamaIndex、Semantic Kernel、AutoGen、CrewAI 或同类框架做过 AI Agent
- 设计并实现过 RAG:Embedding 模型、向量数据库、检索、重排、知识索引
- 用 LoRA、QLoRA、SFT、DPO 等参数高效方法微调过 LLM
- Python 扎实,熟悉 PyTorch、Hugging Face Transformers、PEFT、vLLM 等
- 通过 API 集成过基础模型(OpenAI、Gemini、Anthropic、开源)并评估模型表现
- 熟练掌握 prompt 工程、function calling、结构化输出、工具集成、Agent 编排
- 在生产环境部署过 AI 服务:模型 serving、推理优化、Docker 容器化、云平台
- 了解 AI 系统的可观测性:日志、tracing、监控、评测流水线、性能分析
- 熟悉 Git、CI/CD、API 开发、测试和 Code Review
- 分析与解决问题能力强,能快速评估并采纳新技术;沟通协作好
加分项 / 其他
- 分布式训练与推理优化(DeepSpeed、FSDP、TensorRT-LLM、vLLM、SGLang)
- 多 Agent 系统、工作流编排、规划与记忆架构经验
- 模型评测方法、基准测试、幻觉检测、AI 安全方面的知识
- 做过语音 Agent、对话式 AI 或虚拟助手
- 开源贡献、论文、Kaggle 比赛或个人 AI 项目
- 熟悉 Kubernetes、GPU 基础设施、分布式系统、MLOps 平台
这份 JD 的看点:拿来当「Agent 开发学习地图」最合适:从编排、RAG、工具调用,到评测、可观测、推理优化,全链路都在。注意它把「评测(幻觉检测、A/B、基准)」直接写进了 Agent 工程师的日常职责,说明评测已经不是另一个工种,而是 Agent 工程的一部分。
✅ 你已经有的
- Python 后端、API、Docker、CI/CD 都是你的强项
- LLM 统一接入层(DeepSeek、Dify、LangChain)、SSE 流式输出你在 AI 测试平台里做过
- RAG 知识库模块、MCP 协议有实践(ai-testcase-mcp、mcp-client)
- 自己写过轻量 Agent 框架 nanobot,理解 Agent 循环
⚠️ 差距
- 模型微调(SFT/LoRA/DPO)和 PyTorch/Hugging Face 生态
- 推理服务优化:vLLM、量化、投机解码
- 生产级评测流水线(自动基准 + A/B + 幻觉检测)
- 需要到河内现场办公,不是远程
🎯 怎么补 / 怎么用
- 把你 AI 测试平台里的 RAG 模块补一套评测(见「评测」方向 Dealstitch 那条),这是最容易补、性价比最高的一项
- 跑通一次 LoRA 微调小模型,先有概念即可,不用深入训练
- 用 vLLM 本地部署一个 7B 模型,看看吞吐/延迟怎么测
原帖:https://vn.linkedin.com/jobs/view/agentic-ai-engineer-llm-at-vinfast-4453276050(岗位可能已下线)
查看英文原文(学英语对照用)
VINFAST is a pioneering electric vehicle (EV) company committed to revolutionizing the automotive industry with sustainable and innovative mobility solutions. As a leading player in the EV market, VinFast is dedicated to delivering high-quality, cutting-edge electric vehicles that redefine the driving experience. Our team consists of passionate professionals driven by a shared vision of creating a greener and more sustainable future through innovation, technology, and excellence. In this role, you will be instrumental inSLP Center, using your skills to process and analyze large volumes of data. You will collaborate with diverse teams, including Engineering, Quality Control, VINFAST's suppliers, to createcutting-edgesolutions that will drive the future of transportation. Design, develop, and optimize LLM-based AI Agents for real-world applications with a focus on accuracy, latency, scalability, and reliability. Research, evaluate, and integrate state-of-the-art foundation models (OpenAI, Gemini, Claude, open-source LLMs, etc.) for different AI tasks. Fine-tune, align, and optimize large language models using techniques such as SFT, LoRA/QLoRA, DPO, preference optimization, and model distillation. Design multi-agent systems, agent orchestration, planning, memory, tool calling, and workflow execution architectures. Develop Retrieval-Augmented Generation (RAG) pipelines, including retrieval, reranking, indexing, embedding optimization, and knowledge integration. Design and optimize function calling pipelines for structured task execution and external tool integration. Build evaluation frameworks for AI systems, including automated benchmarking, offline evaluation, online A/B testing, and hallucination detection. Optimize model inference performance through prompt engineering, routing, caching, batching, quantization, speculative decoding, and serving optimization. Develop and maintain APIs and reusable AI services for rapid integration into products. Build observability capabilities including logging, tracing, monitoring, and performance analytics for AI agents in production. Continuously research emerging AI technologies, models, and agent frameworks, rapidly validating and applying promising approaches to production systems. Collaborate with product managers, backend engineers, frontend engineers, and infrastructure teams to deliver production-ready AI features. Produce technical documentation, architecture designs, best practices, and reusable components for the AI platform. Mentor junior engineers, conduct technical reviews, and contribute to engineering standards and AI best practices Requirements Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Software Engineering, or a related field. 3 + years of experience in AI/ML engineering, NLP, LLM applications, or a related software engineering role. (Adjust based on seniority.) Strong understanding of Large Language Models (LLMs), Transformer architectures, and modern generative AI techniques. Hands-on experience building AI agents using frameworks such as LangGraph, LangChain, LlamaIndex, Semantic Kernel, AutoGen, CrewAI, or equivalent. Experience designing and implementing Retrieval-Augmented Generation (RAG) systems, including embedding models, vector databases, retrieval, reranking, and knowledge indexing. Experience fine-tuning LLMs using techniques such as LoRA, QLoRA, Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), or other parameter-efficient training methods. Strong programming skills in Python and familiarity with AI/ML libraries such as PyTorch, Hugging Face Transformers, PEFT, vLLM, or equivalent. Experience integrating foundation models through APIs (OpenAI, Gemini, Anthropic, or open-source LLMs) and evaluating model performance. Solid understanding of prompt engineering, function calling, structured output generation, tool integration, and agent orchestration. Experience deploying AI services in production environments, including model serving, inference optimization, containerization (Docker), and cloud platforms. Knowledge of observability for AI systems, including logging, tracing, monitoring, evaluation pipelines, and performance analysis. Familiarity with software engineering best practices, including Git, CI/CD, API development, testing, and code review. Strong analytical and problem-solving skills with the ability to rapidly evaluate and adopt emerging AI technologies. Excellent communication skills and ability to collaborate effectively in cross-functional teams. Preferred Qualifications Experience with distributed model training and inference optimization (DeepSpeed, FSDP, TensorRT-LLM, vLLM, SGLang, etc.). Experience with multi-agent systems, workflow orchestration, planning, and memory architectures. Knowledge of model evaluation methodologies, benchmarking, hallucination detection, and AI safety. Experience building voice agents, conversational AI, or virtual assistants. Contributions to open-source AI projects, research publications, Kaggle competitions, or personal AI projects are a plus. Familiarity with Kubernetes, GPU infrastructure, distributed systems, and MLOps platforms is preferred. Benefits Competitive salary Premium healthcare package, including PVI insurance & annual health check-ups 13th-month salary & performance bonuses to reward your contributions Enjoy preferential pricing for services within the Vingroup ecosystem including Vinmec, Vinpearl, and Vinschool... Opportunity to collaborate with and learn from industry-leading professionals in the automotive domain Work Location: Technopark Tower, Vinhomes Ocean Park, Gia Lam, Hanoi, Vietnam With respect to all your personal data shared to VinFast in the application and the entire recruitment process of VinFast, by clicking “Apply”, submitting your resumé/CV and/or participating in VinFast's recruitment process, you agree that you have read VinFast's Personal Data Protection Policy ("Policy") posted at https://vinfastauto.com/vn_vi/dieu-khoan-phap-ly or https://vinfast.vn/privacy-policy/ , you agree to the Policy and consent for VinFast to process your personal data in accordance with the Policy and the applicable regulations on personal data protection. To all recruitment agencies : VinFast does not accept agency resumes. Please do not forward resumes to our careers alias or other VinFast employees. VinFast is not responsible for any fees related to unsolicited resumes.
Agent 开发限定国家★★★★★
AI Engineer — Agentic AI Platform
Awear.ai · 美国及部分美洲国家
💰 $120K–$200KAgent 运行时模型路由记忆/RAG工具集成评测
主要职责
- 平台 Agent 运行时:核心编排层,管理 Agent 会话、串联推理步骤、路由到工具、代表用户执行动作
- 多模型供应商集成与路由:在区域模型和任务专用模型之间动态切换,路由时综合考虑延迟、成本、能力
- RAG 架构与记忆检索:让 Agent 能访问用户持久化的加密上下文,并随时间越用越准
- 工具与技能集成流水线:把 Agent 接入外部 API、设备能力和平台自有功能
- Agent 评测与可观测:度量 Agent 质量、暴露失败、让团队看清 Agent 在生产环境的真实表现
- 端侧推理优化:与移动端和固件团队一起判断哪些环节可以在设备本地运行,降低延迟和云依赖
- Prompt 架构与系统设计:结构化提示框架、系统指令、上下文管理模式,保证 Agent 行为一致
- 直接向 CEO 汇报并与全体工程团队协作;写生产代码,为平台上每一次 Agent 交互的质量负责;参与模型选型和路由决策
任职要求
- 4 年以上软件工程经验,至少 2 年直接做过生产环境的 LLM 系统
- 深入的 LLM 集成经验:不只是调 API,而是懂得怎么构建可靠、可扩展、低延迟的 AI 流水线
- RAG 架构经验:向量库、Embedding 流水线、检索优化、上下文窗口管理
- 熟悉 LangChain、LlamaIndex、AutoGen 等 Agent 框架和编排模式,并对它们的优缺点有自己的判断
- Python 工程能力扎实——写的是生产代码,不是研究用 notebook
- 有大规模评测 LLM 输出的经验:建 eval、度量质量、发现回归
- 理解不同 LLM 在能力、延迟、成本、隐私上的取舍,并在生产中做过路由决策
- 积极使用 AI 辅助开发工具(Cursor、Claude、Copilot)并有自己的心得
- 书面英语好,能在远程分布式团队中高效工作
- 必须有美国工作授权
加分项 / 其他
- 端侧/边缘推理经验(Core ML、ONNX、TensorFlow Lite)
- 多 Agent 系统设计
- 常开/流式推理架构
- 隐私保护 AI 系统
- 可穿戴或移动 AI 平台经验
这份 JD 的看点:很典型的创业公司「第一个 AI 工程师」JD:没有细分的岗位边界,Agent 运行时、模型路由、记忆、评测、端侧推理全归你。可以学到 Agent 系统里几个关键设计问题:怎么按延迟/成本/能力做模型路由,怎么把评测和可观测性放进 Agent 层。
✅ 你已经有的
- 模型统一接入层:你做过多 AI 平台(Dify SSE/HTTP、LangChain、DeepSeek)统一接入,和「多供应商集成与路由」是同一类问题
- 多队列多消费者并发模型,对应 Agent 会话并发管理
- RAG 知识库、MCP 工具集成实践
- 用 AI 编码工具的习惯(你本来就在用)
⚠️ 差距
- 美国工作授权是硬门槛,你目前无法满足
- LLM 输出的大规模评测和回归检测
- 端侧推理(Core ML / ONNX)你没有接触
- 「不只是调 API」的可靠性经验需要有可展示的项目
🎯 怎么补 / 怎么用
- 把你自己项目里的模型路由(按任务/成本切模型)单独抽出来做成可演示的小项目,写清楚取舍
- 补 eval:为 Agent 每一步(选工具、检索、最终答案)打分,而不是只看最终结果(见 Dealstitch 条)
原帖:https://wellfound.com/jobs/4185322-ai-engineer-agentic-ai-platform(岗位可能已下线)
查看英文原文(学英语对照用)
We're building the operating system for the next generation of computing — one where AI agents replace apps and your technology finally works for you instead of the other way around. We're a stealth-mode startup with a world-class founding team with deep roots in consumer AI, extended reality, and wearable technology — including founders of some of the most recognizable hardware and software platforms of the last decade. We're backed by strategic partnerships with leading silicon and manufacturing companies, and we're hiring our first AI engineer to build the intelligence layer at the core of the platform. This is a rare opportunity to architect the agent infrastructure of a platform that doesn't exist yet — at the layer where always-on contextual AI meets a wearable form factor for the first time. Additional product details shared under NDA. THE ROLE The backend engineers build the infrastructure. The mobile engineers build the user facing surfaces. You build what runs between them — the agents themselves. As our first AI Engineer you will own the design and implementation of our agent layer — the pipelines, reasoning chains, memory retrieval systems, tool integrations, and orchestration logic that turn raw LLM capability into a platform that genuinely replaces the app paradigm. You will work directly with the CEO and across the full engineering team to make sure the agent experience is as technically rigorous as it is experientially compelling. This is a hands-on engineering role. You will write production code, own the agentic runtime architecture, and be directly accountable for the quality of every agent interaction on the platform. You will also be a key voice in decisions about which models to use, how to route between them, and how to structure the memory and context systems that make our platform smarter over time. We actively use AI development tools across our engineering team — Cursor, Claude, Copilot — and expect engineers who use them seriously as a core part of their workflow. WHAT YOU'LL BUILD The platform agent runtime — the core orchestration layer that manages agent sessions, chains reasoning steps, routes to tools, and executes actions on behalf of users Multi-provider LLM integration and routing — selecting and switching between regional and task-specific language models dynamically, with latency, cost, and capability all factored into routing decisions RAG architecture and memory retrieval — the systems that give agents access to the user's persistent, encrypted context layer and make responses smarter and more relevant over time Tool and skill integration pipelines — the infrastructure that connects agents to external APIs, device capabilities, and first-party platform features Agent evaluation and observability — the frameworks that measure agent quality, surface failures, and give the team visibility into how agents are actually performing in production On-device inference optimization — working with the mobile and firmware teams to identify which parts of the agent pipeline can run locally on device, reducing latency and cloud dependency as the platform evolves toward wearable hardware Prompt architecture and system design — the structured prompting frameworks, system instructions, and context management patterns that govern agent behavior consistently across the platform WHAT WE'RE LOOKING FOR 4+ years of software engineering experience with at least 2 years working directly on LLM-based systems in production Deep hands-on experience with LLM integration — not just API calls but genuine understanding of how to build reliable, scalable, low-latency AI pipelines Strong experience with RAG architectures — vector databases, embedding pipelines, retrieval optimization, context window management Familiarity with agent frameworks and orchestration patterns — LangChain, LlamaIndex, AutoGen, or similar, with a clear point of view on their strengths and limitations Solid Python engineering skills — you are writing production code, not research notebooks Experience evaluating LLM outputs at scale — building evals, measuring quality, detecting regressions Genuine understanding of the tradeoffs between different LLMs — capability, latency, cost, privacy implications — and experience making routing decisions in production systems Active user of AI-assisted development tools with a genuine point of view on how to use them well Strong written English and proven ability to work effectively in a remote and distributed team Must be authorized to work in the US Strong plus: experience with on-device or edge inference — Core ML, ONNX, TensorFlow Lite, or similar; multi-agent system design; always-on or streaming inference architectures; privacy-preserving AI systems; wearable or mobile AI platform experience WHAT WE OFFER Salary: competitive depending on experience Meaningful early-stage equity Full medical, dental, and vision coverage Fully remote with occasional in-person time in Silicon Valley or Paris for key milestones Awear is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Agent 开发限定国家★★★★★
AI Engineer — Hermes Agent
pst.ag · 纯远程
ReAct/Reflexion多 Agent 协作记忆MCPTemporal
主要职责
- 设计、构建并编排能推理、规划和执行复杂工作流的自主 AI Agent——和传统 LLM 聊天机器人不同,这些 Agent 会与动态环境交互、使用工具、与其他 Agent 协作、尽量少依赖人工
- Agent 架构:使用 Hermes Agent 框架;协作模式含 orchestrator-workers、辩论、层级式 swarm
- 记忆:短期/长期/情景记忆,基于向量库和语义缓存
- 推理与规划:ReAct、CoT、ToT、Plan-and-Solve;动态规划、错误恢复、根据反馈重新规划
- 工具使用:function calling、API grounding(数据库、API、RAG、UI 自动化)
- 生产与评测:为任务完成度、效率、安全性做 Agent 评测(不能只看字面相似度);用 LangSmith、Arize、W&B 做链路追踪;优化延迟、token 成本和可靠性
- 集成与工具:对接 CRM、数据库、Slack、浏览器、REST API、代码解释器;构建自定义工具和沙箱环境,安全地执行代码/命令
任职要求
- Python 精通;熟悉 prompt 工程、few-shot、结构化输出(JSON mode、语法约束)
- 在生产或复杂原型中实现过 ReAct、Reflexion、Toolformer 等 Agent 模式
- 向量库(Pinecone、Weaviate、Qdrant)和 RAG 优化(混合检索、重排)经验
- 熟悉工作流引擎(Temporal、Prefect、Airflow),用于人在回路和持久化执行
- 监控 LLM 应用的经验(prompt trace、token 用量、漂移)
- 构建过使用 MCP 的 Agent,完成多步研究、代码分析或数据工程任务
- 熟悉 Hermes Agent(Nous Research 的自我改进 Agent 框架)的安装、配置和运维
- 计算机相关本科;5 年软件工程经验;熟悉规格驱动开发(SDD);了解 BMAD、GitHub Spec Kit 等方法之一
- 做过生产级 Agent 系统(不是 demo 或聊天机器人)
- 理解 LLM 的局限:幻觉、越狱、prompt 注入和失败模式;了解 MCP 发现模式与上下文协商;掌握上下文管理(prompt 缓存、滑动窗口、语义检索、MCP 资源生命周期)
这份 JD 的看点:这份 JD 是「Agent 设计模式词典」:orchestrator-workers、辩论、层级 swarm、ReAct、Reflexion、情景记忆、重规划……每一个词都值得搜一下弄懂。另一个亮点:把 Temporal 这类工作流引擎当作 Agent「持久化执行」的基础设施——Agent 跑很久、会失败、要重试,本质是分布式工作流问题。
✅ 你已经有的
- Temporal 分布式工作流是你的核心强项(稳定性平台:状态持久化、故障恢复、重试,任务成功率 99.5%),正好对上「持久化执行 / 人在回路」
- MCP 协议实践(ai-testcase-mcp、mcp-client)
- Python 生产级后端、RabbitMQ 多消费者
- 自己做过 Agent 框架 nanobot
⚠️ 差距
- Hermes Agent 是一个很新的框架,你需要装起来玩一遍
- 规格驱动开发(SDD)、BMAD 这类方法论你没实践过
- 向量库深度(混合检索、重排)
- 招聘国家名单里没有中国,需要先问清楚能否接受
🎯 怎么补 / 怎么用
- 用 Temporal 写一个「会失败、会重试、可人工确认」的 Agent 流程演示——这是你独一无二的组合,做成开源项目很有说服力
- 搜一遍 ReAct、Reflexion、Plan-and-Solve、orchestrator-workers 这些模式,每个用 50 行代码复现
原帖:https://wellfound.com/jobs/4560710-ai-engineer-hermes-agent(岗位可能已下线)
查看英文原文(学英语对照用)
About the Role: We are seeking a forward-thinking Agentic AI Engineer to design, build, and orchestrate autonomous AI agents capable of reasoning, planning, and executing complex workflows. Unlike traditional LLM-based chatbots, our agents interact with dynamic environments, use tools, collaborate with other agents, and operate with minimal human intervention. Agent Architecture & Development: Framework: Hermes Agent Collaboration: orchestrator-workers, debate, hierarchical swarms Memory: short/long-term + episodic via vector DBs & semantic caching Reasoning & Planning: Techniques: ReAct, CoT, ToT, Plan-and-Solve Dynamic planning, error recovery, replanning from feedback Tool use: function calling, API grounding (DBs, APIs, RAG, UI automation) Production & Evaluation: Eval: agentic evals for task completion, efficiency, safety (not just lexical) Observability: tracing/logging (LangSmith, Arize, W&B) Optimize: latency, token cost, reliability Integration & Tooling: Connect: CRMs, DBs, Slack, browsers, REST APIs, code interpreters Custom tools + sandboxed envs for safe code/shell execution Technical Skills: Programming: Expert in Python Strong understanding of prompt engineering, few-shot learning, and structured output generation (JSON mode, grammars). Reasoning Patterns: Proven experience implementing agentic patterns (ReAct, Reflexion, Toolformer) in production or complex prototypes. Memory & Retrieval: Experience with vector databases (Pinecone, Weaviate, Qdrant) and RAG optimization (hybrid search, reranking). Orchestration: Familiarity with workflow engines (Temporal, Prefect, Airflow) for human-in-the-loop and durable execution. Observability: Experience monitoring LLM applications (prompt traces, token usage, drift). Model Context Protocol: Built agents that use MCP for multi-step research, code analysis, or data engineering tasks. Agentic Framework : Practical experience with Hermes Agent Education & Experience: Bachelor’s degree in Computer Science, Software Engineering, AI, or related discipline 5 years in software engineering Strong background on Spec-Driven Development ( SDD ) methodology Practical experience installing, configuring, and operating Hermes Agent (the self-improving AI agent framework from Nous Research) Experience building production-grade agentic systems (not just demos or chatbots). Must be well versed with any of the following Method: BMAD ( Breakthrough Method for Agile AI-Driven Development ) Github Spec Kit OpenSec Strong understanding of LLM limitations: hallucinations, jailbreaks, prompt injection, and failure modes. Good understanding of MCP discovery patterns and context negotiation. Strong knowledge of context management in LLM applications: prompt caching, sliding window, semantic retrieval, MCP resource lifecycle.
Agent 开发限定国家★★★★★
Senior AI Engineer(多 Agent IT 运维平台)
Mapgenesys · 美国芝加哥
💰 $80K–$150K多 AgentAIOps渐进式自主护栏混合检索/重排
主要职责
- 提升 Agent 推理:设计和打磨多步推理流水线,让 Agent 根据结果动态规划、执行、重新规划;目标是 80% 以上故障自主解决
- 构建智能知识检索:在企业知识库上实现高级 RAG(混合检索、重排、分块策略、自适应检索)
- 开发持续学习能力:做模式挖掘,发现重复出现的故障模式、构建因果模型、从历史处理记录里自动生成运维手册(runbook)
- 设计多 Agent 协作:Agent 之间如何共享上下文、协商优先级、按故障复杂度和置信度动态分派
- 实现安全与护栏:扩展安全层,加入风险评分、爆炸半径估算、渐进式自主模型(影子模式 → 有人监督 → 完全自主)
- 优化成本与延迟:智能模型路由(复杂推理用强模型,分类用轻量模型)、响应缓存、token 优化
任职要求
- 5 年以上 Python,含 async/await、FastAPI 或类似框架,软件工程基础扎实
- LLM 应用开发:做过 LLM 会推理、规划、用工具的生产系统,prompt 工程是基本功
- 熟悉 LangChain、LangGraph、CrewAI、AutoGen 等 Agent 框架,理解状态管理、checkpoint、条件路由
- RAG 实现经验:向量库、Embedding、混合检索、检索评测——知道朴素 RAG 为什么会失败以及怎么修
- 能在不确定中工作:自己设计、构建、测试、迭代,不等详细需求
- 可观测与生产就绪:Langfuse / OpenTelemetry 链路追踪、结构化日志、prompt/版本追踪、token 用量监控、模型漂移检测
- 成本与性能优化:模型选择权衡(延迟 vs 推理深度)、缓存、批处理、工具调用效率、上下文窗口优化
- 安全与治理意识:prompt 注入防护、数据隔离、密钥处理、PII 脱敏、安全的工具访问模式
加分项 / 其他
- IT 运维、NOC、ITSM 平台经验(ServiceNow、PagerDuty)
- 因果推断或运维数据模式挖掘
- MCP 或类似工具使用框架
- 向量库(pgvector、Pinecone、Weaviate、Qdrant)
- 自主系统的 AI 安全与护栏
- 云原生部署(AWS/GCP/Azure、Kubernetes)
这份 JD 的看点:最值得学的是「渐进式自主」:影子模式 → 有人监督 → 完全自主,这是 Agent 上线的通用安全路径,也是评测和护栏的落点。另外「爆炸半径估算」「风险评分」把 Agent 的动作当成有风险的操作来管理,思路很工程化。
✅ 你已经有的
- 运维平台经验:环境管理平台(K8s、织云、TKEX 统一调度)、CI/CD 插件,你就是 IT 运维/DevOps 的行内人
- opsCopilot(LLM 运维助手)、cle_agent(环境管理智能体)——你的个人项目和这个岗位方向几乎重合
- Python async、FastAPI、Celery 都是主力
- Temporal 的重试/补偿思路可迁移到 Agent 的重新规划
⚠️ 差距
- LangGraph 的状态 checkpoint、条件路由需要上手
- Langfuse / OpenTelemetry 的 LLM 链路追踪
- 混合检索 + 重排的实操
- 只招芝加哥,无法远程到中国
🎯 怎么补 / 怎么用
- 把 opsCopilot 升级成「有护栏的运维 Agent」:加风险评分、只读/低风险/高风险三档权限,并写文章讲清楚——这就是这份 JD 想要的样子
- 用 LangGraph 复现一次「告警 → 诊断 → 建议 → 人工确认」流程
原帖:https://wellfound.com/jobs/4612622-sr-ai-engineer(岗位可能已下线)
查看英文原文(学英语对照用)
Senior AI Engineer Location: Chicago, US Type: Full-time Department: AI & Automation — IT Operations ──────────────────────────────────────────────────────────── About the Role We're building an AI-powered autonomous operations platform that uses multi-agent AI to detect, diagnose, and resolve IT infrastructure incidents — with minimal human intervention. Our platform coordinates specialized AI agents that collaborate across the incident lifecycle, from initial alert through resolution. We have a working MVP processing real incidents end-to-end. We need someone to take the AI layer from functional to intelligent — improving agent reasoning, building smarter retrieval, and making the system genuinely learn from every incident it handles. You'll own the AI layer: agent reasoning, multi-step planning, tool use, knowledge retrieval, and continuous learning. ──────────────────────────────────────────────────────────── What You'll Do • Advance agent reasoning — Design and refine multi-step reasoning pipelines where AI agents dynamically plan, execute, and re-plan based on results. Push toward 80%+ autonomous incident resolution • Build intelligent knowledge retrieval — Implement advanced RAG patterns including hybrid search, re-ranking, chunking strategies, and adaptive retrieval across enterprise knowledge bases • Develop continuous learning capabilities — Build pattern mining that discovers recurring incident patterns, constructs causal models, and auto-generates operational runbooks from historical resolutions • Design multi-agent collaboration — Architect how agents share context, negotiate priorities, and dynamically route work based on incident complexity and confidence levels • Implement safety and guardrails — Extend the safety layer with risk scoring, blast radius estimation, and progressive autonomy models (shadow → supervised → autonomous) • Optimize cost and latency — Implement intelligent model routing (powerful models for complex reasoning, lightweight models for classification), response caching, and token optimization strategies ──────────────────────────────────────────────────────────── What You Bring Must Have • 5+ Python — async/await, FastAPI or similar frameworks, strong software engineering fundamentals • Hands-on LLM application development — You've built production systems where LLMs reason, plan, and use tools. Prompt engineering is second nature • Experience with agentic AI frameworks — LangChain, LangGraph, CrewAI, AutoGen, or similar. You understand state management, checkpointing, and conditional routing in agent workflows • RAG implementation experience — Vector databases, embedding models, hybrid search, retrieval evaluation. You know why naive RAG fails and how to fix it • Comfortable with ambiguity — You'll design, build, test, and iterate in a fast-moving environment. You figure out the right approach, not wait for detailed specs • Observability & production readiness — Tracing (Langfuse/OpenTelemetry), structured logging, prompt/version tracking, monitoring token usage, model drift detection • Cost-performance optimization — Model selection tradeoffs (latency vs reasoning depth), caching strategies, batching, tool-call efficiency, and context window optimization • Security & governance awareness — Prompt injection mitigation, data isolation, secrets handling, PII redaction, and secure tool access patterns Nice to Have • Experience with IT operations, NOC, or ITSM platforms (ServiceNow, PagerDuty, etc.) • Knowledge of causal inference or pattern mining from operational data • Experience with MCP (Model Context Protocol) or similar tool-use frameworks • Familiarity with vector databases (pgvector, Pinecone, Weaviate, Qdrant) • Prior work on AI safety and guardrails for autonomous systems • Experience deploying LLM systems in cloud-native environments (AWS/GCP/Azure, Kubernetes) ──────────────────────────────────────────────────────────── Why Join • Greenfield AI architecture — You're building autonomous agents that reason, learn, and collaborate — not maintaining legacy ML pipelines • Real production impact — Every improvement directly reduces mean time to resolution and operational toil • Cutting-edge applied AI — Multi-agent coordination, agentic RAG, continuous learning from operations data • Ownership — Small team, high autonomy. You'll shape the AI architecture, not just implement tickets
Agent 开发限定国家★★★★★
AI Agent & Automation Developer
Cherry Bekaert · 美国
💰 $117,400–$172,500Copilot StudioSemantic KernelAzure AI上下文工程企业自动化
主要职责
- 用 LLM、规划算法和决策框架设计并实现 AI Agent
- 开发支持自主性、交互性和任务完成的 Agent 架构
- 把 Agent 集成到应用、API 或工作流里(聊天机器人、自动化工具等)
- 与研究员、工程师、产品团队协作,迭代 Agent 能力
- 通过反馈回路、强化学习或用户交互优化 Agent 行为
- 做「上下文工程」,保证 Agent 在对的时间拿到对的信息:prompt 工程(把信息注入 prompt)、微调(把知识训练进模型)、RAG(文档向量化检索)、工具使用(子 Agent 或确定性模型补充上下文)、以及混合方案
- 监控性能、做评测、实现安全与护栏机制
- 维护 Agent 逻辑、设计决策和依赖的文档
- 与业务方和工程团队协作,把 Agent 集成到微软平台:Copilot Studio、Azure AI、Semantic Kernel SDK;用微软的编排工具构建企业级 Copilot,打通 Microsoft 365、Dynamics 和自定义业务应用
任职要求
- 计算机、AI、机器学习或相关专业本科
- 有 Agent 系统或智能自动化经验,同类岗位至少 2 年
- Python(或类似)编程能力强,熟悉 LangChain、OpenAI API、Hugging Face、PyTorch 等库
- 理解基于 Agent 的建模、多 Agent 系统或强化学习
- 有用 LLM/生成式 AI 模型构建应用的经验
- 熟悉 API 开发、后端服务和部署流水线
- 能在实验性、快节奏的环境中独立工作
- 沟通和人际能力强,能与业务及技术方建立关系;具备分析思维和数据驱动决策能力;理解数据质量、完整性和系统集成
这份 JD 的看点:它对「上下文工程」的拆解很清楚:prompt 注入、微调、RAG、工具使用、混合方案,五种把信息交给模型的方式,可以当作一张面试答题表。另外能看到传统企业买的是微软那一套,和创业公司自研框架是完全不同的路线。
✅ 你已经有的
- Python、API、后端服务、部署流水线都满足
- LLM 应用(Dify、LangChain)和 RAG 有实践
⚠️ 差距
- 微软全家桶(Copilot Studio、Azure AI、Semantic Kernel)你没有接触
- 岗位要求美国工作资格且需到办公室,无法满足
- 偏业务咨询,技术成长性一般
🎯 怎么补 / 怎么用
- 如果只想了解企业 Agent 的另一条路线,花半天读 Semantic Kernel 文档即可,不必深入
原帖:https://www.linkedin.com/jobs/view/ai-agent-automation-developer-at-cherry-bekaert-4422343399(岗位可能已下线)
查看英文原文(学英语对照用)
Ranked among the largest accounting and consulting firms in the country and consistently recognized as a Great Place to Work, Cherry Bekaert delivers innovative advisory, assurance and tax services to our clients. We are proud to foster a collaborative environment focused on enabling your career growth and continuous professional development. Our Business Transformation team is looking for an AI Agent & Automation Developer. The AI Agent & Automation Developer will have the opportunity to work in a hybrid model from any of Cherry Bekaert's US office locations. We are seeking a talented and motivated AI Agent & Automation Developer to join our team. The ideal candidate will have hands-on experience building autonomous AI agents using large language models (LLMs), agent frameworks, and modern AI infrastructure. As an AI Agent & Automation Developer, you will design, implement, and deploy intelligent agents that operate autonomously within digital environments, driving innovation and operational excellence. This role supports organizational transformation by solving complex business challenges through intelligent automation, agent orchestration, and advanced analytics. You must be adaptable, collaborative, and eager to learn new technologies to address multifaceted business problems. A successful candidate is a proactive problem-solver with strong software engineering skills, ready to thrive in a fast-paced environment and communicate effectively with both technical and business stakeholders. As An AI Agent & Automation Developer, You Will Design and implement AI agents using LLMs, planning algorithms, and decision-making frameworks. Develop agent architectures that support autonomy, interactivity, and task completion. Integrate agents into applications, APIs, or workflows (e.g., chatbots, automation tools). Collaborate with researchers, engineers, and product teams to iterate on agent capabilities. Optimize agent behavior through feedback loops, reinforcement learning, or user interaction. Engineer and manage context for agents to ensure they access the right information at the right time. This includes applying methods such as: Prompt Engineering – injecting information directly into prompts. Fine-tuning – training models on additional data to embed knowledge (where applicable). Retrieval-Augmented Generation (RAG) – embedding documents into vector databases and retrieving semantically relevant information. Tool Use – leveraging sub-agents or deterministic models to locate and add relevant context. Hybrid Approaches – combining two or more methods to maximize reliability and relevance. Monitor performance, conduct evaluations, and implement safety and guardrail mechanisms. Maintain thorough documentation of agent logic, design decisions, and dependencies. Stay up to date with industry developments and advancements in AI, machine learning, and agent-based technologies. Collaborate with stakeholders and engineering teams to integrate AI agents with Microsoft-specific platforms such as Copilot Studio, Azure AI, and Semantic Kernel SDK. Build and deploy enterprise-grade Copilots using Microsoft’s orchestration tools, enabling seamless interaction across Microsoft 365, Dynamics, and custom business applications. What You Bring To The Role Bachelor’s degree in Computer Science, AI, Machine Learning, or a related field. Proven experience in agent-based systems, or intelligent automation, with a minimum of 2 years in a similar role. Strong programming skills in Python (or similar) and experience with AI/ML libraries (e.g., LangChain, OpenAI API, Hugging Face, PyTorch). Understanding of agent-based modeling, multi-agent systems, or reinforcement learning. Experience building applications with LLMs or generative AI models. Familiarity with API development, backend services, and deployment pipelines. Ability to work independently in experimental and fast-paced environments. Excellent communication and interpersonal skills, with the ability to build strong relationships with business and technical stakeholders. Analytical mindset with the ability to solve complex problems and make data-driven decisions. Demonstrated understanding of data quality, integrity, and systems integration. What You Can Expect From Us Our shared values that foster inclusion and belonging including uncompromising integrity, collaboration, trust, and mutual respect The opportunity to innovate and do work that motivates and engages you A collaborative environment focused on enabling you to further your career growth and continuous professional development Competitive compensation and a total rewards package that focuses on all aspects of your wellbeing Flexibility to do impactful work and the time to enjoy your life outside of work Opportunities to connect and learn from professionals from different backgrounds and with different cultures Benefits Information Cherry Bekaert cares about our people. We offer competitive compensation packages based on performance that recognize the value our people bring to our clients and our Firm. The salary range for this position is included below. Individual salaries within this range are determined by a variety of factors including but not limited to the role, function and associated responsibilities, a candidate’s work experience, education, knowledge, skills, and geographic location. In addition, we offer a comprehensive, high-quality benefits program which includes annual bonus, medical, dental, and vision care; disability and life insurance; generous Paid Time Off; retirement plans; Paid Care Leave; and other programs that are dedicated to enhancing your personal and work life and providing you and your family with a measure of financial protection. Pay Range $117,400-$172,500 About Cherry Bekaert Cherry Bekaert, ranked among the largest assurance, tax and advisory firms in the U.S., serves clients across industries in all 50 U.S. states and internationally. For more details, visit https://www.cbh.com/disclosure/ Cherry Bekaert provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, national origin, citizenship status, protected veteran status, disability status, or any other category protected by applicable federal, state or local laws. https://careers.cbh.com/legal-disclosures/ contains further information regarding the firm's compliance with federal, state and local recruitment and hiring laws. This role is expected to accept applications for at least five calendar days and may continue to be posted until a qualified applicant is selected or the position has been cancelled. Candidates must demonstrate eligibility to work in the United States. Cherry Bekaert will not provide work sponsorship for this position. Cherry Bekaert LLP and Cherry Bekaert Advisory LLC are members of Allinial Global, an accountancy and business advisory global association. Visit us at https://careers.cbh.com/ and follow us on LinkedIn, Instagram, Twitter and Facebook. © 2026 Cherry Bekaert. All Rights Reserved.
Agent 开发限定国家★★★★★
Senior AI Engineer(临床 AI Agent)
Cadence · 美国
💰 基础薪资 $180,000–$245,000临床 AgentLLM-as-judgeCI/CD 评测RAG护栏
主要职责
- 设计、构建、部署临床 AI Agent:对患者上下文推理、调用工具、生成照护建议
- 负责大规模 LLM 工作流的可靠性、可观测性和成本效率
- 在临床知识库、治疗方案和实时患者数据上构建并优化 RAG 流水线
- 开发评测框架:离线基准、安全测试、回归测试集,以及接入 CI/CD 的 LLM-as-judge 流水线
- 设计多步 Agent 编排:规划、记忆、工具使用、错误恢复、人在回路的升级路径
- 与临床、产品、工程团队协作,把患者照护需求转成 AI 系统设计
- 跟进 AI 与工程最佳实践,持续提高质量、性能和架构水平
任职要求
- 计算机、工程或相关专业本科/硕士,或同等工作经验
- 5 年以上软件工程经验,其中 2 年以上构建生产环境的 AI/ML 系统
- 在高速成长的环境里有从设计到上线的端到端所有权
- 熟悉 LLM API(OpenAI、Anthropic、开源模型):prompt 工程、工具调用/function calling、结构化输出
- 有 RAG 系统经验:Embedding、向量库、检索优化、grounding
加分项 / 其他
- 熟悉 Agent 框架或编排模式(工具调用、planner、多 Agent 协作)
- 有领域微调经验(SFT、RLHF、LoRA)
- 医疗或强监管行业经验(HIPAA、SOC 2、临床数据处理)
这份 JD 的看点:这份 JD 是「评测怎么落地」的好样本:离线基准 + 安全测试 + 回归集 + LLM-as-judge,并且接进 CI/CD。可靠性直接影响患者的领域,评测不是加分项而是门禁。整组样本里薪资也是最明确、最高的(基础 $180–245K)。
✅ 你已经有的
- CI/CD 流水线是你的主场(50+ 业务线接入,插件调用过亿次),把「评测接进 CI」对你来说是熟悉的工程动作
- RAG、LLM 接入、工具调用(MCP)都有实践
- 质量部背景,天然懂回归测试
⚠️ 差距
- 医疗/强监管领域(HIPAA、SOC 2)
- LLM-as-judge 的校准与实战
- 美国地区,未说明接受境外
🎯 怎么补 / 怎么用
- 用你的 CI 经验做一个「LLM 评测门禁」小项目:每次改 prompt 自动跑回归集,指标掉了就阻断合并
原帖:https://www.linkedin.com/jobs/view/4463106321/(岗位可能已下线)
查看英文原文(学英语对照用)
Sixty million Medicare seniors live with chronic disease. The care system sees most of them twice a year. Cadence is building the infrastructure to support them every day. Cadence is a clinical AI company that delivers continuous, proactive care for older adults with chronic conditions like hypertension, heart failure, and diabetes. We pair patients with a dedicated clinical team, integrate deeply into health system EMRs and workflows, and use our Clinical Intelligence platform to monitor vitals, surface risk early, optimize medications, and close care gaps between visits. The result: patients engage with care 100x more than before Cadence, clinicians focus on judgment instead of administrative work, and Medicare saves $2M a week. We operate as a full clinical care delivery organization, not a software vendor. Our clinicians work alongside health system partners, extending the reach of local primary care providers into patients' homes. We're now applying AI agents across these workflows – from alert review and medication titration to lifestyle coaching and care coordination – with clinicians always in control of clinical decisions. The Role We're hiring a Senior AI Engineer to design and ship production agents that synthesize real-time clinical data, surface proactive care recommendations, and take action on behalf of clinicians. You will take agent architectures from prototype to production in a domain where reliability directly affects patient outcomes. You'll own the full lifecycle of AI-powered clinical workflows: retrieval, reasoning, tool use, evaluation, and safety guardrails. If you've built agentic systems at scale and want your work to matter beyond engagement metrics, we’d love to speak with you. What You'll Do Design, build, and deploy clinical AI agents that reason over patient context, invoke tools, and generate care recommendations. Own reliability, observability, and cost efficiency of LLM-powered workflows at scale. Build and optimize RAG pipelines over clinical knowledge bases, treatment protocols, and real-time patient data. Develop evaluation frameworks: offline benchmarks, safety tests, regression suites, and LLM-as-judge pipelines wired into CI/CD. Design multi-step agent orchestration: planning, memory, tool use, error recovery, and human-in-the-loop escalation paths. Collaborate with clinical, product, and engineering teams to translate patient care needs into AI system design. Stay at the forefront of AI and engineering best practices, continuously pushing the team to raise the bar on quality, performance, and architecture. What You Need Bachelor's or Master’s degree in Computer Science, Engineering or related field, or equivalent work experience 5+ years of software engineering experience, with 2+ years building AI/ML-powered systems in production Experience in a high-growth, fast-paced environment with end-to-end ownership from design through production Hands-on experience with LLM APIs (OpenAI, Anthropic, open-source models) including prompt engineering, tool use / function calling, and structured outputs Experience building RAG systems: embeddings, vector stores, retrieval optimization, and grounding Experience with agent frameworks or orchestration patterns (tool calling, planners, multi-agent coordination) preferred Fine-tuning experience (SFT, RLHF, LoRA) on domain-specific tasks preferred Healthcare or regulated-industry experience (HIPAA, SOC 2, clinical data handling) preferred Compensation Our job titles may span more than one career level. The base salary for this role typically ranges between $180,000 – $245,000 annual base salary, depending on experience, skills, seniority, and business needs. In addition to base salary, this role is eligible for equity as part of the total compensation package. Actual compensation may vary by location. Benefits & Perks Competitive pay & equity* Fully remote Comprehensive health coverage: Medical, dental & vision Paid time off 401k plan + matching Paid parental leave Home office stipend benefit offerings may vary depending on job profile, job level and worker type Cadence is committed to equal opportunity and fairness regardless of race, color, religion, sex, gender identity, sexual orientation, nation of origin, ancestry, age, physical or mental disability, country of citizenship, medical condition, marital or domestic partner status, family status, family care status, military or veteran status or any other basis protected by local, state or federal laws. A notice to Cadence applicants: Our Talent team only directs candidates to apply through our official careers page at https://www.cadence.care/our-team. Cadence will never refer you to external websites, ask for payment or personal information, or conduct interviews via messaging apps. We receive all applications through our website and anyone suggesting otherwise is not with Cadence. If you require a reasonable accommodation during the interview or hiring process, please notify your recruiter. … 更多 职位发布中说明的福利 401(k) 订阅相似职位 Senior AI Engineer,美国 关 在招人? 立即发布职位 关于 无障碍模式 人才解决方案 社区准则 招聘专版 营销解决方案 隐私政策和条款 广告设置 广告 销售解决方案 移动端 小型企业 安全中心 LinkedIn Corporation © 2026 年 需要咨询? 欢迎访问帮助中心。 管理帐号和隐私 前往设置。 推荐透明度 详细了解推荐内容。 选择语言 العربية (阿拉伯语) বাংলা (孟加拉语) Čeština (捷克语) Dansk (丹麦语) Deutsch (德语) Ελληνικά (希腊语) English (英语) Español (西班牙语) فارسی (波斯语) Suomi (芬兰语) Français (法语) हिंदी (印地语) Magyar (匈牙利语) Bahasa Indonesia (印尼语) Italiano (意大利语) עברית (希伯来语) 日本語 (日语) 한국어 (韩语) मराठी (马拉地语) Bahasa Malaysia (马来语) Nederlands (荷兰语) Norsk (挪威语) ਪੰਜਾਬੀ (旁遮普语) Polski (波兰语) Português (葡萄牙语) Română (罗马尼亚语) Русский (俄语) Svenska (瑞典语) తెలుగు (泰卢固语) ภาษาไทย (泰语) Tagalog (他加禄语) Türkçe (土耳其语) Українська (乌克兰语) Tiếng Việt (越南语) 简体中文 (简体中文) 正體中文 (繁体中文)
Agent 开发需到岗海外★★★★★
Principal AI Agent Development Engineer / Architect
Bybit · 马来西亚 · 吉隆坡
Agent SDKMCP 矩阵ADLCGo/Python评测体系
主要职责
- Agent SDK 与运行时:设计并实现编排引擎、运行时和核心原语;保证 Agent 循环既可控、可验证,又保持对话性和适应性
- MCP 矩阵扩展:把内部系统包装成 MCP server / 子 Agent / Agent 技能;统一工具发现、动态注册、RBAC、限流和版本管理
- Agent 开发生命周期(ADLC):搭建适应 Agent 非确定性行为的工具链,覆盖完整闭环:规格 → 评测 → 构建 → 测试 → A/B → 灰度 → 监控 → 迭代
- 对效果与成本端到端负责:负责准确率、解决率、成本、延迟等指标;通过数据反馈建立 Agent 个性化优化闭环
任职要求
- 计算机/软件工程等相关专业本科及以上
- P7:6 年以上经验(含 3 年以上 AI/Agent);P7+:8 年以上,且有 0 到 1 做 Agent 平台的经验
- 熟练使用 Go 或 Python,能做生产级系统设计并撰写 TRD 文档
- LLM 应用工程:精通 LLM API、prompt 工程、ReAct 模式、上下文窗口管理(token 预算、历史压缩、子 Agent 压缩)
- 生产级 Agent 经验:完整的 0 到 1 上线并持续迭代——明确区别于 Demo/POC
- 熟悉至少一种编排框架(LangGraph / Dify / Eino / Embabel / OpenAI Agents SDK 等),有多 Agent 协作实战(Agent Teams / 并行 / DAG / Handoff)
- 至少精通一项相邻能力:RAG 工程(向量检索、混合检索、重排)或评测体系(LLM-as-Judge、Golden Set、回归测试、A/B)
- 工程思维:能量化效果与成本(token 优化思路);理解 PII 脱敏和 prompt 注入防护;能识别金融场景的合规边界
加分项 / 其他
- 0 到 1 搭过 Agent 仿真/基准测试平台(τ-Bench 规模)
- 私有化 LLM 部署:vLLM / Triton / DeepSeek / Qwen / GLM
- 领域经验:加密 AI(风控/KYC/合规)、AIOps(告警根因/诊断工具)、客服 AI、代码 AI
- AI 安全研究:prompt 注入、越狱防御、智能合约 AI 安全等方面的成果
- 对 LangChain / LangGraph / Dify / MCP / Eino / Vercel AI SDK 等的开源贡献
- 公开的技术影响力:博客、演讲、论文、开源项目
这份 JD 的看点:最值得读的是 ADLC(Agent 开发生命周期):规格 → 评测 → 构建 → 测试 → A/B → 灰度 → 监控 → 迭代。这就是「怎么把 Agent 当成软件来工程化」的一张流程图,评测是里面一个正式环节。另外「把内部系统统一包成 MCP server,并做 RBAC/限流/版本管理」是企业落地 MCP 的真实形态。
✅ 你已经有的
- Go + Python 双语言,正好是 JD 要求的「Go 或 Python」
- MCP 实践(ai-testcase-mcp、mcp-client);CI/CD 与灰度发布是你的主场(50+ 业务线、发布门禁)
- 多 Agent / 工作流编排:Temporal 的 DAG 与调度经验可迁移
- 质量部背景,和「评测 → A/B → 灰度」这条链路天然亲近
⚠️ 差距
- 「0 到 1 上线的生产级 Agent 平台」需要有可展示的案例
- 评测体系(LLM-as-Judge、Golden Set)的实战
- Agent 仿真/基准平台、vLLM 私有化部署
- 地点在吉隆坡,要到岗
🎯 怎么补 / 怎么用
- 把「ADLC」当作你的项目框架:给自己的 Agent 项目补上规格文档 → 评测集 → 灰度开关 → 监控看板,写成一篇完整的复盘
原帖:https://job-boards.eu.greenhouse.io/bybit/jobs/4905494101(岗位可能已下线)
查看英文原文(学英语对照用)
About Us Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 — connecting users to the future of digital finance. Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution. Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services. Key Responsibilities Agent SDK & Runtime — Design and implement the orchestration engine, runtime, and core primitives; ensure the agentic loop is both steerable and verifiable while remaining conversational and adaptive. MCP Matrix Expansion — Wrap internal systems as MCP servers / sub-agents / agent skills; unify tool discovery, dynamic registration, RBAC, rate limiting, and version management. Agent Development Lifecycle (ADLC) — Build toolchains adapted to the non-deterministic behavior of agents, covering the full loop: Spec → Eval → Build → Test → A/B → Canary → Monitoring → Iteration. End-to-End Ownership of Effectiveness & Cost — Own accuracy, resolution rate, cost, and latency metrics; build agent personalization optimization loops through data feedback. Requirements Education & Experience: Bachelor's degree or above in CS/SE or related field; P7: 6+ years (including 3+ years in AI/Agent); P7+: 8+ years with 0-to-1 Agent platform experience. Engineering Foundation: Proficient in Go or Python; capable of production-grade system design and TRD documentation. LLM Application Engineering: Deep expertise in LLM APIs, Prompt Engineering, ReAct patterns, and context window management (token budgeting / history compression / sub-agent compression). Production-Grade Agent Experience: Full 0-to-1 launch with continuous iteration loop — clearly distinguished from Demo/POC work. Agent Engineering Practice: Familiar with at least one orchestration framework (LangGraph / Dify / Eino / Embabel / OpenAI Agents SDK, etc.); hands-on experience with Multi-Agent collaboration (Agent Teams / parallel / DAG / Handoff). Adjacent Agent Capabilities: Deep expertise in at least one of: RAG engineering (vector retrieval, Hybrid Search, Reranking) or evaluation systems (LLM-as-Judge, Golden Set, regression testing, A/B). Engineering Mindset: Ability to quantify effectiveness and cost (token optimization thinking); understanding of PII desensitization and prompt injection protection; ability to identify compliance boundaries in financial scenarios. Bonus Points 0-to-1 experience building an Agent simulation / benchmarking platform (τ-Bench scale) Private LLM deployment: vLLM / Triton / DeepSeek / Qwen / GLM Domain expertise: Crypto AI (risk control / KYC / compliance), AIOps (alert RCA / diagnostic tooling), customer service AI, code AI AI security research: published work on prompt injection, jailbreak defense, Smart Contract AI security, etc. Open-source contributions to LangChain / LangGraph / Dify / MCP / Eino / Vercel AI SDK or similar Public technical influence: blog posts / conference talks / papers / OSS projects Why Join Us At Bybit, we are committed to fostering a supportive and enriching work environment. Our benefits include: - Study Growth Fund: We support your professional development and continuous learning. - Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation. - Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world. - Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company. - Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.
Agent 开发需到岗海外★★★★★
Principal AI Engineer, AI Agent Development
OKX · 香港 / 新加坡
多 Agent 系统强化学习上下文工程交易/风控 Agent团队领导
主要职责
- 主导为加密交易所运营量身定制的自主 AI Agent 的设计、开发和全生命周期管理(客户助理、交易机器人、风险检测 Agent、反欺诈 Agent 等)
- 架构多 Agent 系统(MAS)和框架,支持目标导向、推理、规划和基于强化学习的 Agent
- 与产品、工程、业务、运营团队合作,定义 Agent 角色和可衡量的 KPI
- 推动实验、仿真和持续学习流水线,优化 Agent 表现
- 跟进 AGI、LLM 集成 Agent、认知架构的前沿,把合适的创新用到生产系统
- 指导并带领 AI 科学家和工程师团队,参与制定 AI 路线图和研究策略
- 保证 AI 方案在加密交易所领域的可扩展性、安全性和合规性
任职要求
- 计算机、AI、机器学习、机器人等相关专业硕士或博士,且有 10 年以上行业经验
- 在真实环境中构建并部署过自主 AI Agent;理想情况下:设计过多 Agent 框架或生态(AutoGPT、OpenAgents、LangGraph、LangFuse、Dify、Coze 等)
- 有目标导向 Agent、上下文工程、记忆管理、工具使用和基于强化学习的决策经验
- 开发过实时决策场景的 Agent,如交易、游戏、客服或运营
- 在快节奏、高规模环境(最好是金融、加密或高频领域)端到端交付 AI 系统的成绩
加分项 / 其他
- 顶级 AI 实验室或科技巨头经验
- 高频交易、金融科技、Web3 或加密平台经验
- 创业公司或研究密集型环境背景,能适应不确定性和速度
这份 JD 的看点:这是「Agent 方向的天花板岗位」的样子:不只写代码,而要定义 Agent 角色和 KPI、带团队、定研究路线。适合用来了解 Agent 场景在金融里能落到哪些具体产品(客服、交易、风控、反欺诈),而不是拿来当投递目标。
✅ 你已经有的
- Dify / LangChain 等框架接触,Agent 框架自研经验(nanobot)
⚠️ 差距
- 10 年以上 + 硕博门槛,你 9 年 / 本科,差一截
- 强化学习和多 Agent 研究背景
- 领导 AI 科学家团队的经历
- 香港/新加坡到岗
🎯 怎么补 / 怎么用
- 把它当作「远景 JD」:看清 5 年后 Agent 高级岗要求的是定义问题和带人,而不只是写代码
原帖:https://job-boards.greenhouse.io/okx/jobs/6690158003(岗位可能已下线)
查看英文原文(学英语对照用)
OKX will be prioritising applicants who have a current right to work in Singapore, and do not require OKX's sponsorship of a visa. Who We Are At OKX, we believe that the future will be reshaped by crypto, and ultimately contribute to every individual's freedom. OKX is a leading crypto exchange, and the developer of OKX Wallet, giving millions access to crypto trading and decentralized crypto applications (dApps). OKX is also a trusted brand by hundreds of large institutions seeking access to crypto markets. We are safe and reliable, backed by our Proof of Reserves. Across our multiple offices globally, we are united by our core principles: We Before Me , Do the Right Thing , and Get Things Done . These shared values drive our culture, shape our processes, and foster a friendly, rewarding, and diverse environment for every OK-er. OKX is part of OKG, a group that brings the value of Blockchain to users around the world, through our leading products OKX, OKX Wallet, OKLink and more. About the Opportunity We are seeking a Principal Engineer with a deep expertise in autonomous AI agent architecture and deployment, to spearhead the design, development, and optimization of intelligent agent systems on our global crypto exchange platform. This is a senior role requiring a hands-on AI science background combined with strategic oversight, technical direction, and cross-functional collaboration. You will work at the cutting edge of AI and crypto, designing AI agents that can autonomously reason, act, and evolve in complex, real-time environments such as trading, security, compliance, and customer interaction. What You’ll Be Doing Lead the design, development, and lifecycle management of autonomous AI agents tailored to crypto exchange operations (e.g., customer assistant, trading bots, risk detection agents, fraud prevention agents, etc.). Architect multi-agent systems (MAS) and frameworks to support goal-directed, reasoning, planning, and reinforcement-learning-based agents. Collaborate with product, engineering, business, and operation teams to define agent roles and measurable KPIs. Drive experimentation, simulations, and continuous learning pipelines for agent performance optimization. Stay ahead of the curve on developments in AGI, LLM-integrated agents, and cognitive architectures; apply relevant innovations to production systems. Mentor and lead a team of AI scientists and engineers; help define the AI roadmap and research strategy. Ensure system scalability, security, and regulatory compliance of AI-powered solutions in the crypto exchange domain. What We Look For In You PhD or Master's in Computer Science, AI, Machine Learning, Robotics or related field with at least 10 years of industry experience . Proven experience building and deploying autonomous AI agents in live environments. Ideal candidates will have: Designed multi-agent frameworks or agent ecosystems (e.g., AutoGPT, OpenAgents, LangGraph, LangFuse, Dify, Coze, etc.). Experience with goal-oriented agents, context engineering, memory management, tool use, and RL-based decision making. Developed agents for real-time decision-making, such as trading, gaming, customer support, or operations. Track record of delivering end-to-end AI systems in fast-paced, high-scale environments (preferably finance, crypto, or high-frequency domains). Nice to Haves Experience in top-tier AI labs or Tech industry leaders Experience in high-frequency trading, fintech, Web3, or crypto platforms is a strong advantage Startups or research-intensive environments preferred; must be comfortable operating with ambiguity and speed Perks & Benefits Competitive total compensation package L&D programs and Education subsidy for employees' growth and development Various team building programs and company events Wellness and meal allowances Comprehensive healthcare schemes for employees and dependants More that we love to tell you along the process! Notice: All official OKX vacancies are published on this website. While roles may appear on selected third-party platforms from time to time, information on other sites may be inaccurate or outdated. If in doubt, please apply directly through our official careers website. Information collected and processed as part of the recruitment process of any job application you choose to submit is subject to OKX 's Candidate Privacy Notice .
Agent 开发需到岗海外★★★★★
Principal Engineer, Agent Infrastructure & Memory Architecture
OKX · 香港 / 新加坡
长期记忆知识图谱GraphRAGContext GraphAgent 基础设施
主要职责
- 为 AI Agent 设计并实现长期记忆基础设施:可审计、有版本、可回滚的中间件,支撑持久的、有状态的推理
- 设计并负责端到端的 GraphRAG 与知识图谱流水线:从数据摄取、schema 设计到图索引、检索优化,以及与 LLM 推理系统的集成
- 构建可扩展的上下文管理基础设施:把上下文从 prompt 级技巧提升为 Context Graph 这样的结构化、可维护系统,支持长时程和跨会话推理
- 把前沿 Agent 研究转化成生产级系统:原型验证新架构,建立记忆、检索和基础设施可靠性的最佳实践
- 保证 Agent 基础设施的可靠性、可观测性和治理:可追溯的记忆操作、评测框架、防止漂移和退化的保护措施
任职要求
- 计算机、AI、机器学习、机器人等相关专业硕士或博士,10 年以上行业经验
- 在长期记忆、知识图谱、上下文管理三个领域中,至少精通一个
加分项 / 其他
- 顶级 AI 实验室或科技巨头经验
- 高频交易、金融科技、Web3、加密平台经验
- 创业公司/研究密集环境,适应不确定性和速度
这份 JD 的看点:要求很少,但每个词都很重:「可审计、有版本、可回滚的记忆中间件」,意思是把 Agent 的记忆当成数据库来治理,而不是把聊天记录塞进向量库。适合用来理解 Agent 记忆这个方向的工程深度。
✅ 你已经有的
- 中间件与后端基础设施的工程背景(K8s、分布式任务、持久化)
- RAG 知识库模块的实践
⚠️ 差距
- 知识图谱、GraphRAG 是全新的领域
- 10 年以上 + 硕博门槛
- 香港/新加坡到岗
🎯 怎么补 / 怎么用
- 如果对记忆感兴趣,先用 Neo4j + LLM 做一个「从文档抽实体关系并问答」的小 GraphRAG 实验,弄清它和普通 RAG 的差别
原帖:https://job-boards.greenhouse.io/okx/jobs/7648794003(岗位可能已下线)
查看英文原文(学英语对照用)
OKX will be prioritising applicants who have a current right to work in Singapore, and do not require OKX's sponsorship of a visa. Who We Are At OKX, we believe that the future will be reshaped by crypto, and ultimately contribute to every individual's freedom. OKX is a leading crypto exchange, and the developer of OKX Wallet, giving millions access to crypto trading and decentralized crypto applications (dApps). OKX is also a trusted brand by hundreds of large institutions seeking access to crypto markets. We are safe and reliable, backed by our Proof of Reserves. Across our multiple offices globally, we are united by our core principles: We Before Me , Do the Right Thing , and Get Things Done . These shared values drive our culture, shape our processes, and foster a friendly, rewarding, and diverse environment for every OK-er. OKX is part of OKG, a group that brings the value of Blockchain to users around the world, through our leading products OKX, OKX Wallet, OKLink and more. About the Opportunity We are seeking an AI Infrastructure Engineer to design and build the foundational systems that power next-generation AI agents. This role sits at the intersection of research and production engineering, focusing on scalable long-term memory , knowledge graph infrastructure , and robust context management systems . You will transform cutting-edge research concepts (e.g., Context Graphs, GraphRAG, agent memory architectures) into auditable, maintainable, and production-grade infrastructure that enables reliable, stateful AI systems. What You’ll Be Doing Architect and implement long-term memory infrastructure for AI agents , including auditable, versioned, and rollback-capable middleware that supports persistent, stateful reasoning. Design and own end-to-end GraphRAG and knowledge graph pipelines , from ingestion and schema design to graph indexing, retrieval optimization, and integration with LLM-based reasoning systems. Build scalable context management infrastructure , elevating context from prompt-level techniques to structured, maintainable systems such as Context Graphs that enable long-horizon and cross-session reasoning. Translate advanced agent research into production-grade systems , prototyping emerging architectures and establishing best practices for memory, retrieval, and infrastructure reliability. Ensure reliability, observability, and governance of agent infrastructure , including traceable memory operations, evaluation frameworks, and safeguards against drift and degradation. What We Look For In You PhD or Master's in Computer Science, AI, Machine Learning, Robotics or related field with at least 10 years of industry experience . Possess a deep expertise in at least one of the following areas: long-term memory , knowledge graphs , or context management . Nice to Haves Experience in top-tier AI labs or Tech industry leaders Experience in high-frequency trading, fintech, Web3, or crypto platforms is a strong advantage Startups or research-intensive environments preferred; must be comfortable operating with ambiguity and speed Perks & Benefits Competitive total compensation package L&D programs and Education subsidy for employees' growth and development Various team building programs and company events Wellness and meal allowances Comprehensive healthcare schemes for employees and dependants More that we love to tell you along the process! Notice: All official OKX vacancies are published on this website. While roles may appear on selected third-party platforms from time to time, information on other sites may be inaccurate or outdated. If in doubt, please apply directly through our official careers website. Information collected and processed as part of the recruitment process of any job application you choose to submit is subject to OKX 's Candidate Privacy Notice .
Agent 开发地区开放★★★★★
Python Backend Development Talent with RAG and Agentic AI Experience
Toptal(客户项目) · 仅限亚洲
Python/FastAPIRAGLangGraph微服务Docker/K8s
主要职责
- 设计可扩展 RAG 系统和 AI 聊天机器人的端到端架构
- 用 FastAPI 等框架开发 Python 后端服务、REST API 和微服务
- 构建文档摄取、分块、Embedding、索引、检索、重排流水线
- 用 Azure AI Search 或同类向量数据库实现向量检索
- 设计能规划、推理、使用工具和做决策的自主 AI Agent
- 用 LangGraph、LangChain、CrewAI、AutoGen 等编排框架开发多步 AI 工作流
- 集成商业与开源 LLM:Azure OpenAI、OpenAI、Anthropic Claude、Gemini 等
- 实现对话记忆、会话管理、上下文管理和 Agent 协作模式
- 把 AI 工作流与 API、数据库、企业系统、外部工具连接起来
- 开发能处理并发 AI 负载的异步高性能服务
- 实现 prompt 管理、结构化输出、护栏、回退逻辑和模型评测流程
- 建立日志、监控、追踪、可观测、安全与错误处理规范;用 Docker、Kubernetes、Azure 或 AWS 容器化部署 AI 服务
- 把业务需求转成技术设计、交付里程碑和生产级 AI 方案;在团队中独立负责架构和实现决策
任职要求
- 8 年以上 Python 后端开发经验
- 设计 REST API、微服务、异步服务和分布式后端系统的扎实经验
- 构建过生产级 RAG 应用;理解 Embedding、文档分块、语义搜索、向量索引、检索策略和重排
- 开发过 AI Agent 和多步 LLM 工作流,熟悉 LangGraph、LangChain、CrewAI、AutoGen 等 Agent 框架
- 通过 OpenAI、Azure OpenAI、Anthropic Claude、Gemini 或开源模型 API 集成 LLM
- 能设计超越基础 prompt 工程的 AI 架构;能把 AI 应用与 API、数据库、数据管道、企业系统集成
- 为生产服务实现安全、监控、日志、追踪和可观测
- 用 Docker 和 Azure/AWS 等云平台部署容器化应用
- 能独立把业务需求转化为可扩展的技术方案;在分布式环境下沟通协作能力强
- 可与美国工作时间有数小时重叠
这份 JD 的看点:和你的技术栈重合度最高,且限定亚洲,是这组样本里最现实的一个。「架构超越基础 prompt 工程」这句话很关键:它要的是能把 RAG 和 Agent 做成稳定后端服务的人,而不是调 prompt 的人。
✅ 你已经有的
- Python 9 年,FastAPI/Django/Celery、微服务、异步、RabbitMQ 都是你的主场
- Docker/K8s、监控、CI/CD 生产经验
- RAG 知识库(AI 测试平台)、LangChain、LLM 多模型接入
- 分布式系统:Temporal、gRPC、千级并发任务
⚠️ 差距
- LangGraph / CrewAI 的实战案例
- Azure AI Search / Azure OpenAI 这套云服务
- 需要一个「生产级 RAG + Agent 后端」的可展示项目(你的项目多为内网/私有)
- 要与美国时区重叠数小时(对你就是晚上工作)
🎯 怎么补 / 怎么用
- 把 AI 测试平台里可公开的部分抽成开源示例:FastAPI + RAG + LangGraph + 追踪,README 写清架构;这是投这类岗位最有力的材料
- 熟悉 Azure OpenAI 的接入方式(和 OpenAI API 差别不大,主要是鉴权与部署名)
查看英文原文(学英语对照用)
Headquarters: Summary: We are seeking an experienced Python Backend Developer to design, build, and deploy scalable AI-powered applications using Retrieval-Augmented Generation, large language models, and agentic AI frameworks. The role will focus on delivering a production-grade RAG system and AI chatbot that can securely integrate with enterprise data, APIs, databases, and cloud services. General information: The organization is developing an AI-powered platform and requires an experienced individual contributor to build its RAG architecture and conversational AI capabilities. The developer will work closely with a distributed team and should be available for several hours of overlap with US working hours. The project involves designing AI systems that go beyond basic prompt engineering, including multi-step workflows, autonomous agents, vector search, knowledge retrieval, memory management, and tool integration. The solution must be scalable, secure, observable, and suitable for production use. The technology environment includes Python, FastAPI, large language models, LangGraph, LangChain, vector databases, Azure AI services, AWS, Docker, Kubernetes, and microservice-based architectures. Task and deliverables: Design the end-to-end architecture for a scalable RAG system and AI chatbot. Develop Python backend services, REST APIs, and microservices using FastAPI or similar frameworks. Build document ingestion, chunking, embedding, indexing, retrieval, and reranking pipelines. Implement vector search solutions using Azure AI Search or comparable vector databases. Design autonomous AI agents capable of planning, reasoning, tool usage, and decision-making. Develop multi-step AI workflows using LangGraph, LangChain, CrewAI, AutoGen, or similar orchestration frameworks. Integrate commercial and open-source LLMs, including Azure OpenAI, OpenAI, Anthropic Claude, Gemini, and comparable models. Implement conversation memory, session management, context management, and agent collaboration patterns. Connect AI workflows with APIs, databases, enterprise systems, and external tools. Develop asynchronous, high-performance services capable of handling concurrent AI workloads. Implement prompt management, structured outputs, guardrails, fallback logic, and model evaluation processes. Establish logging, monitoring, tracing, observability, security, and error-handling standards. Containerize and deploy AI services using Docker, Kubernetes, Azure, or AWS. Translate business requirements into technical designs, delivery milestones, and production-ready AI solutions. Collaborate with the wider team while independently owning architecture and implementation decisions. Required experience: Required: 8 or more years of professional Python backend development experience. Required: Strong experience designing REST APIs, microservices, asynchronous services, and distributed backend systems. Required: Hands-on experience building production-grade RAG applications. Required: Strong understanding of embeddings, document chunking, semantic search, vector indexing, retrieval strategies, and reranking. Required: Hands-on experience developing AI agents and multi-step LLM workflows. Required: Experience with agentic AI frameworks such as LangGraph, LangChain, CrewAI, AutoGen, or comparable platforms. Required: Experience integrating LLMs through OpenAI, Azure OpenAI, Anthropic Claude, Gemini, or open-source model APIs. Required: Ability to design AI architecture beyond basic prompt engineering. Required: Experience integrating AI applications with APIs, databases, data pipelines, and enterprise systems. Required: Experience implementing security, monitoring, logging, tracing, and observability for production services. Required: Experience deploying containerized applications using Docker and cloud platforms such as Azure or AWS. Required: Ability to independently translate business requirements into scalable technical solutions. Required: Strong communication and collaboration skills in a distributed working environment. Required: Availability for several hours of overlap with US working hours. To apply: https://weworkremotely.com/remote-jobs/toptal-python-backend-development-talent-with-rag-and-agentic-ai-experience
Agent 开发地点待确认★★★★★
Backend Engineer, AI (Agent Systems)
ActAI · 新加坡
推理编排低延迟流式输出生产运维Python/Node
主要职责
- 构建并运营为 AI 功能提供服务的生产后端系统
- 围绕模型设计推理流水线、编排层和服务边界
- 负责生产问题:监控、日志、告警、故障响应
- 优化推理、缓存、批处理和流式输出的延迟与吞吐
任职要求
- 扎实的生产环境后端工程基础
- 运行高吞吐、低延迟服务的经验
- 熟悉 AI 推理模式(LLM、Embedding、多模态)
- 能在高负载下调试分布式系统
- 倾向于快速上线,并从生产表现中学习
加分项 / 其他
- 技术栈:Python、NodeJS、PyTorch、OpenAI / Anthropic / 开源 LLM、SQL 与 NoSQL、Kubernetes、Docker
这份 JD 的看点:JD 很短,但描述的是 Agent 产品里最容易被忽视的一层:模型之上、用户之下的编排与推理服务。公司强调「长时间运行的工作流、持久上下文、真实任务完成」,这本质是可靠性工程。「结果导向」写法(Outcomes)也值得学:稳定、清晰的 API、故障快速定位。
✅ 你已经有的
- Python、Node/TS、Docker、K8s 全部是你的日常
- SSE 流式推送、多队列并发、异步任务是你做过的
- 生产运维:稳定性平台、故障恢复、监控
- Temporal 对「长时间运行的工作流」有对应经验
⚠️ 差距
- PyTorch 与推理层细节(批处理、缓存策略)
- 未写明薪资与年限,需要沟通
- 新加坡地点,是否接受中国远程要确认
🎯 怎么补 / 怎么用
- 选一个开源模型,自己写一个带流式、缓存、限流的推理服务,测一下延迟和吞吐,写成压测报告
原帖:https://www.linkedin.com/jobs/view/4469122710/(岗位可能已下线)
查看英文原文(学英语对照用)
About ActAI There are over 5 billion users using basic applications today such email, notes, tasks, calendar and they're not AI-native. Our mission is to build proactive applications for anyone in the world, who are not used to complex prompting. We aim to bring intelligence to conversations, errands, organising and workflows, with minimal to no prompting. Our product focuses on achieving high reliability for long-running workflows, persistent context, and real-world task completion. We believe products will greatly reduce hallucinations. Our objective is to organise anyone's life, allowing us all to spend time on valuable and meaningful things. Role As a Backend Engineer, AI, you own the inference and orchestration layer that powers every AI interaction in the product. Your work sits between models and users, where latency, correctness, reliability, and cost directly impact real-world experience. You will build and operate production systems that turn model capability into fast, stable, observable APIs used across mobile and desktop clients. Focus Build and operate backend systems that serve AI-powered features in production. Design inference pipelines, orchestration layers, and service boundaries around models. Own production concerns: monitoring, logging, alerting, and incident response. Optimize latency and throughput across inference, caching, batching, and streaming. Ideal Experiences Strong backend engineering fundamentals in production environments. Experience running high-throughput, low-latency services. Familiarity with AI inference patterns (LLMs, embeddings, multimodal). Comfortable debugging distributed systems under load. Bias toward shipping and learning from production behavior. Outcomes Backend systems run reliably at scale, handling production AI traffic with low latency and high throughput. APIs are stable, clear, and support seamless integration with frontend and ML systems. Production incidents are quickly detected, diagnosed, and resolved, minimizing user impact. Iterative improvements based on real usage continuously increase system performance and reliability. Tech Stack Python NodeJs Pytorch OpenAI / Anthropic / open-source LLMs SQl & noSQL Kubernetes Docker How We Work The best products today in the world were built by small, world class teams. We are a high talent density and hands-on team. We make decisions collectively, move at rapid speed, striking a balance between shipping high quality work and learning. Joining our team requires the ability to bring structure, exercise judgment, and execute independently. Our goal is to put in hands of our users a truly magical product Interview process If there appears to be a fit, we'll reach to schedule 3, but no more than 4 interviews. Applications are evaluated by our technical team members. Interviews will be conducted via virtual meetings and/or onsite. We value transparency and efficiency, so expect a prompt decision. If you've demonstrated the exceptional skills and mindset we're looking for, we'll extend an offer to join us. This isn't just a job offer; it's an invitation to be part of a team that's bringing AI to have practical benefits to billions globally.
Agent 开发需到岗海外★★★★★
Senior AI Engineer(前线交付方向)
H2O.ai · 新加坡
Agent 交付断网部署vLLM 容量规划K8s/HelmLLMOps
主要职责
- 设计并构建 Agentic AI 系统和多 Agent 框架,为企业客户自动化复杂多步工作流
- 用 RAG、微调、prompt 工程、function calling 和工具使用开发 LLM 应用;在更合适的场景构建传统 ML/深度学习方案(预测、分类、表格模型)
- 实现护栏、人在回路检查点和评测框架,确保负责任的生产级 AI 行为
- 端到端负责开发全周期:问题定义、数据探索与清洗,到集成、测试和生产部署
- 构建向企业应用暴露 AI 能力的后端服务和 API
- 在客户环境(云、本地、完全断网)部署和运维模型与服务,包括基础设施、Kubernetes、Helm、Ingress/TLS 和 GPU 启用
- 合理规划部署规模:对推理服务做压测,测出真实吞吐、延迟、并发,转成客户可预算的 GPU 规模建议
- 建设 LLMOps 基础设施,支持生产环境的持续监控与改进
- 客户对接:作为面向客户的窗口,处理技术支持工单,做培训和研讨会,参与售前和 POC,规划并跟踪多个并行的客户项目
任职要求
- 3 年以上动手工程经验,其中 1 年以上直接做 Agentic AI 系统,有端到端 AI 应用上线的记录
- 构建过 LLM 应用:RAG 流水线、Agent 工作流、微调模型等
- 在云或企业环境(AWS、Azure、GCP、本地 K8s)部署模型和 AI 服务的经验
- 有安全认证环境或完全断网环境的经验:离线交付、镜像同步、私有仓库、依赖打包
- 深入理解 GenAI 与 Agentic AI:prompt 工程、RAG、微调、模型评测、护栏、LLMOps 和 MCP;了解传统机器学习和评测指标
- Python 工程能力强;熟悉 PyTorch/TensorFlow/scikit-learn 与 LangChain/LlamaIndex 等工具
- 能用 vLLM 等在单机与多机上服务开源权重模型,理解张量并行、最大长度、显存利用率、并发等参数
- 理解推理性能取舍:首 token 时间、token 间延迟、吞吐、并发、上下文长度
- 熟悉 AWS 基础(EC2、GPU 实例、VPC、IAM、S3、CloudWatch)和 GPU 成本优化;Kubernetes 与 Helm 熟练
- 后端开发:REST API、容器化、AI 应用的 CI/CD;LLM 可观测(Fluentd、Prometheus、Grafana、Langfuse)
- 能与非技术人员沟通复杂 AI 系统;能写运维手册、容量说明、故障复盘等文档
加分项 / 其他
- 公共部门、金融、医疗等强监管行业的 AI 部署经验
- 前线/客户现场工程师经验
- 紧跟 AI 领域最新进展
这份 JD 的看点:这是「Forward Deployed / 前线工程师」的典型 JD,可以看到 Agent 落地到企业客户时的另一半工作:断网环境交付、GPU 容量规划、培训与售前。它说明 Agent 岗位不只是写代码,也包括部署与客户成功。
✅ 你已经有的
- K8s、Docker、Nginx、CI/CD,环境管理平台就是做「环境交付」的
- Python 后端、API、LLM 应用
- CLI 工具、运维手册类文档习惯
⚠️ 差距
- vLLM 与 GPU 推理服务的容量规划
- 断网环境交付、Helm 经验
- 客户对接、培训、售前(沟通量大)
- 新加坡到岗
🎯 怎么补 / 怎么用
- 如果对「前线工程」感兴趣:在本地用 vLLM 起一个模型,压测并写一份「这个模型要多少 GPU 撑多少并发」的容量说明
原帖:https://sg.linkedin.com/jobs/view/ai-engineer-at-h2o-ai-4452417146(岗位可能已下线)
查看英文原文(学英语对照用)
Founded in 2012, H2O.ai is on a mission to democratize AI. As the world’s leading agentic AI company, H2O.ai converges Generative and Predictive AI to help enterprises and public sector agencies develop purpose-built GenAI applications on their private data. With a focus on Sovereign AI—secure, compliant, and infrastructure-flexible deployments—H2O.ai delivers solutions that align with the highest standards of data privacy and control. Our open-source technology is trusted by over 20,000 organizations worldwide, including more than half of the Fortune 500. H2O.ai powers AI transformation for companies like AT&T, Commonwealth Bank of Australia, Chipotle, Workday, Progressive Insurance, and NIH. H2O.ai partners include NVIDIA, Dell Technologies, Deloitte, Ernst & Young (EY), Snowflake, AWS, Google Cloud Platform (GCP), VAST Data and MinIO. H2O.ai’s AI for Good program supports nonprofit groups, foundations, and communities in advancing education, healthcare, and environmental conservation. With a vibrant community of 2 million data scientists worldwide, H2O.ai aims to co-create valuable AI applications for all users. H2O.ai has raised 256 million from investors, including Commonwealth Bank, NVIDIA, Goldman Sachs, Wells Fargo, Capital One, Nexus Ventures and New York Life. For more information, visit www.h2o.ai . About This Opportunity We are looking for a Senior AI Engineer who builds things that matter. You will design and ship end-to-end AI solutions for some of APAC's most complex enterprise problems - spanning agentic AI systems, LLM applications, and production ML pipelines. This is a hands-on engineering role embedded within a customer-facing field team, meaning your work will be seen, used and evaluated by real enterprises from day one. You will work alongside Kaggle Grandmasters, ML engineers, and domain experts to deliver AI that goes beyond demos - into production, into workflows, and into measurable business outcomes. This position is based in Singapore. What You Will Do AI, Agentic AI & LLM Engineering Design and build agentic AI systems and multi-agent frameworks that automate complex, multi-step workflows for enterprise customers. Develop LLM-powered applications using RAG, fine-tuning, prompt engineering, function calling and tool use. Build and deploy classical ML and deep learning solutions (e.g. forecasting, classification, tabular models) where they outperform or complement LLM-based approaches. Implement guardrails, human-in-the-loop checkpoints, and evaluation frameworks to ensure responsible, production-grade AI behavior. End-to-End AI Application Development Own the full development lifecycle from problem framing, data exploration, cleaning through integration, testing and production deployment. Build scalable backend services and APIs that expose AI capabilities to enterprise applications. Deploy and operate models and services in customer environments, including cloud, on-prem, and fully air-gapped, covering infrastructure provisioning, Kubernetes, Helm, ingress/TLS and GPU enablement. Right-size deployments: load-test serving stacks to measure real throughput, latency and concurrency, then turn those results into concrete GPU sizing guidance customers can plan and budget against. Build LLMOps infrastructure for continuous monitoring and improvement in production. Customer Engagement & Delivery Serve as the customer-facing point of contact across a range of stakeholder profiles, including data scientists, engineers and business leaders, translating real-world problems into AI solutions. Triage and resolve technical customer support tickets from initial diagnosis through to resolution. Run hands-on training, workshops and enablement sessions to help customer teams get the most out of the platform. Collaborate closely with internal stakeholders, including Product Managers, Engineering, and Kaggle Grandmasters to deliver cohesive, high-quality solutions. Contribute to pre-sales and proof-of-concept engagements: building fast, credible demonstrations that win technical trust, and scoping engagements realistically. Solid project management skills: able to plan, prioritize and track multiple concurrent customer engagements to completion. End-to-end ownership of what is shipped into a customer environment. What We Are Looking For Experience & Background 3+ years of hands-on engineering experience, including 1+ years working directly on Agentic AI systems with a track record of end-to-end AI application development and deploying applications into production. Demonstrable experience building LLM-powered applications: RAG pipelines, agentic workflows, fine-tuned models or similar. Experience deploying models and AI services in cloud or enterprise environments (AWS, Azure, GCP, on-prem Kubernetes). Experience working in security-accredited or fully air-gapped environments with no internet egress: offline delivery, image mirroring, private registries and dependency bundling. Skills & Capabilities Deep understanding of modern GenAI and Agentic AI concepts: prompt engineering RAG, fine-tuning, model evaluation, guardrails, LLMOps and MCP. Good understanding of traditional machine learning deep learning, feature engineering and evaluation metrics. Strong Python engineering skills; experience with ML frameworks (PyTorch, TensorFlow, scikit-learn) and LLM tooling (LangChain, LlamaIndex, or equivalent). Able to serve open-weight models with vLLM (or equivalent) across single and multi-node setups, sizing deployments from a model config and tuning key flags (tensor parallelism, max model length, GPU memory utilization, concurrency). Fluent in serving-performance trade-offs: time-to-first-token, inter-token latency, throughput, concurrency and context length. Working knowledge of AWS fundamentals (EC2, GPU instances, VPC, IAM, S3, CloudWatch) and GPU cost optimization. Kubernetes and Helm fluency: deploying, upgrading, and debugging workloads. Backend development skills: REST APIs, containerization (Docker/Kubernetes), and CI/CD pipelines for AI applications. Experience with observability for LLM systems: log aggregation and monitoring (e.g. Fluentd, Prometheus, Grafana) and LLM-specific tracing tooling (e.g. Langfuse or equivalent). Strong problem-solving instincts: comfortable with ambiguity, able to move fast without sacrificing engineering quality. Comfortable operating within change-control or accreditation processes where fixes cannot be pushed freely. Clear communicator who can explain complex AI systems and technical constraints to non-technical stakeholders. Able to write good, consistent technical documentation, both internal (runbooks, sizing notes, post-incident write-ups) and customer-facing. How to Stand Out From the Crowd Experience in public sector, financial services, healthcare or other regulated industry AI deployments. Prior experience in a customer-facing or forward deployed engineering role. Keeps up with the latest advancements in the AI and agentic AI space and is able to incorporate them into client deliverables. Working knowledge of hardening practices relevant to enterprise/public-sector deployments: mTLS, log masking and redaction, secrets management, RBAC and audit requirements. Kaggle or competitive ML experience. Why H2O.ai? Market leader in total rewards Remote-friendly culture Flexible working environment Be part of a world-class team Career growth H2O.ai is committed to creating a diverse and inclusive culture. All qualified applicants will receive consideration for employment without regard to their race, ethnicity, religion, gender, sexual orientation, age, disability status or any other legally protected basis. H2O.ai is an innovative AI cloud platform company, leading the mission to democratize AI for everyone. Thousands of organizations from all over the world have used our cutting-edge technology across a variety of industries. We’ve made it easy for people at all levels to generate breakthrough solutions to complex business problems and advance the discovery of new ideas and revenue streams. We push the boundaries of what is possible with artificial intelligence. H2O.ai employs the world’s top Kaggle Grandmasters, the community of best-in-the-world machine learning practitioners and data scientists. A strong AI for Good ethos and responsible AI drive the company’s purpose. Please visit www.H2O.ai to learn more.
Agent 开发需到岗海外★★★★★
AI Engineer
Distyl AI · 英国 · 伦敦
Python/TSAgent 工作流评测框架客户驱动内部 LLM 平台
主要职责
- 构建生产级 AI 系统:用 LLM 设计、开发并部署强壮的 AI 应用,包括 prompt 工程、Agent 工作流、工具使用和全栈 AI 产品
- 直接与客户合作:和企业干系人一起理解复杂问题,转成有影响力的 AI 方案
- 主导系统架构:设计可扩展的生产 AI 系统架构,兼顾性能、可靠性、成本和可维护性
- 开发内部平台:参与内部 LLM 应用平台 Distillery,构建跨客户部署复用的基础设施、工具和工作流
- 严格评测 AI 系统:开发评测框架,度量模型在准确率、延迟、成本、可靠性、安全性上的表现
- 交付生产级系统:保证可观测、可靠、安全、可维护
- 提高工程水准:改进开发流程、评测实践和部署策略
任职要求
- 3 年以上专业软件工程经验
- 精通 Python 或 TypeScript
- 在生产中构建并部署过 LLM 应用或 AI Agent
- 熟悉 LangChain、LlamaIndex、Guardrails、MCP 或 Agent 框架等现代 LLM 工具
- 实现过 RAG 流水线、工具使用或多步 AI 工作流
- 理解 AI 系统评测、调试和可观测
- 用现代 DevOps 实践构建可靠的生产系统
加分项 / 其他
- 企业环境部署 AI 系统的经验
- 跨云平台(AWS、GCP、Azure)经验
- Agent 架构与长时程任务执行经验
- 负责任 AI 实践:可审计性与治理
这份 JD 的看点:JD 结构清爽,可以对照看「要求」里的必备项和加分项怎么分层:Python/TS + LLM 应用 + RAG/工具 + 评测/可观测是必备;企业部署、多云、长时程 Agent、负责任 AI 是加分。3 年经验的门槛,说明 Agent 岗对「已上过生产」比对「年限」更看重。
✅ 你已经有的
- Python / TypeScript 都是你的强项(React 18 + TS ★5)
- Agent 与 LLM 应用、MCP、RAG 有实践
- 可观测、DevOps 是你的主场
⚠️ 差距
- 评测框架的系统经验
- 客户现场协作、方案沟通
- 伦敦到岗,需要签证
🎯 怎么补 / 怎么用
- 把「评测 + 可观测」补成你的第二项拿手技能:给 Agent 项目接一套 Langfuse 追踪,加十几条回归用例
原帖:https://uk.linkedin.com/jobs/view/ai-engineer-at-distyl-4385939295(岗位可能已下线)
查看英文原文(学英语对照用)
About Distyl AI Distyl is an applied AI technology company partnering with the world’s most ambitious institutions to rearchitect critical operations for the frontier of AI. Our customers include the largest companies in telecom, healthcare, insurance, manufacturing, consumer goods, and global social organizations. We research and deploy technologies that power AI-native operations — both for our partners and for Distyl itself. Our work spans research into self-constructing systems, the development of the most reliable execution of AI systems, and products that transform mission-critical workflows. As a result, Distyl's technologies affect some of the world's largest operations — from hundreds of millions of consumer interactions to tens of millions of supply chain transactions and millions of patient journeys. Distyl is backed by leading investors including Lightspeed Venture Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s. What We’re Looking For We’re opening an office in London, UK and looking for AI Engineers to design and deploy production-grade AI systems powered by LLMs. At Distyl, AI Engineers work directly with Fortune 500 companies to transform complex workflows using cutting-edge AI. You’ll build and ship real-world AI applications — from intelligent agents to full-stack AI products — and see them operate at scale in mission-critical environments. This role is highly hands-on. You’ll collaborate closely with customers, define system architectures, and build reliable, high-impact AI systems from prototype to production. Engineers at Distyl also help shape technical direction across major customer engagements, guiding enterprise teams through AI adoption and deployment. Key Responsibilities Build Production AI Systems: Design, develop, and deploy robust AI applications using LLMs, including prompt engineering, agent workflows, tool use, and full-stack AI products Work Directly with Customers: Partner closely with enterprise stakeholders to understand complex problems and translate them into impactful AI solutions Lead System Architecture: Design scalable architectures for production AI systems, balancing performance, reliability, cost, and maintainability Develop Our Internal Platform: Contribute to Distillery, our internal LLM application platform, by building reusable infrastructure, tools, and workflows used across customer deployments Evaluate AI Systems Rigorously: Develop evaluation frameworks that measure model performance across accuracy, latency, cost, reliability, and safety Ship Production-Grade Systems: Ensure systems meet high standards for observability, reliability, security, and maintainability Raise the Engineering Bar: Improve development workflows, evaluation practices, and deployment strategies as our AI platform continues to evolve Who You Are 3+ years of professional software engineering experience Strong proficiency in Python or TypeScript Experience building and deploying LLM-powered applications or AI agents in production Experience with modern LLM tooling such as LangChain, LlamaIndex, Guardrails, MCP, or agent frameworks Experience implementing RAG pipelines, tool use, or multi-step AI workflows Strong understanding of AI system evaluation, debugging, and observability Experience building reliable production systems with modern DevOps practices Experience deploying AI systems in enterprise environments is a plus Experience working across cloud platforms (AWS, GCP, or Azure) is a plus Experience with agent architectures and long-horizon task execution is a plus Familiarity with responsible AI practices, including auditability and governance is a plus What We Offer Competitive salary, meaningful equity, and a comprehensive benefits package Workplace Pension Scheme with employer contributions Private Medical Insurance (PMI) offered Flexible Time Off + Holidays Lunch provided on office days Access to state-of-the-art models, generous usage of modern AI tools, and real-world business problems Ownership of high-impact projects across top enterprises A mission-driven, fast-moving culture that prizes curiosity, pragmatism, and excellence We believe diverse perspectives make our work stronger and more impactful. We are an equal opportunity employer and evaluate all applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, or any other legally protected characteristic. We encourage candidates from all backgrounds to apply.
🧪 AI 评测(14)
从传统 QA 转型的「AI 质量工程」,到评测平台、RL 环境、红队,再到前沿实验室的研究岗。
AI 评测需到岗海外★★★★★
GenAI / Agentic AI Evaluation Engineer(质量、安全与可靠性)
EY(安永)GDS Assurance · 印度 · 诺伊达
评测策略LLM-as-judge评测 harnessOWASP LLM Top 10红队
主要职责
- 定义并落地 GenAI 系统的评测策略,覆盖问答助手、摘要、信息抽取、起草、Agent 和多步工作流
- 把业务用例翻译成结构化评测计划:范围、假设、成功标准、数据集、指标、红队场景、阈值和报告要求
- 推动标准化:可复用的评测模板、测试用例库、评分 rubric 和报告格式
- 为产品团队设计结构化的数据集需求,确保覆盖:核心用户旅程与主要业务意图;边界情况(罕见 prompt、歧义问题、上下文缺失);对抗样例(恶意 prompt、越狱、prompt 注入);偏见与公平样例
- 给出数据集是否充分的指引和统计覆盖方法(最小样本量、分布均衡、场景矩阵、按意图/风险分层)
- 构建可复用评测流水线:答案质量(正确、相关、完整、清晰);依据与忠实度(RAG:faithfulness、context precision/recall、幻觉率、引用质量);Agent 行为(工具调用准确率、工具误用、目标完成度、步骤正确性、多余动作、循环检测、工具输出安全);运营质量(延迟、成本/token 预算、吞吐、稳定性、重试、故障恢复)
- 以「校准」的方式结合 LLM-as-judge 与人工评测(rubric 设计、抽样方案、一致性检查)
- 用 Python 实现自动化评测 harness:批量跑场景集、指标可配置、可复现(run ID 与产物)、存储 trace 和输出以便审计
- 参照 OWASP Top 10 for LLM Applications 做结构化红队:prompt 注入(直接 + 间接)与工具劫持;敏感数据/PII 泄露;不安全的输出处理;训练数据泄露/记忆探测;模型拒绝服务/「钱包耗尽」攻击
- 把评测集成进开发生命周期:发布前回归门禁、CI 检查、跨模型版本/prompt/工具/检索器的基准对比
- 对 Agent 工作流做对抗测试:工具误用/权限过大、未授权动作执行、经工具/连接器外泄、通过检索文档注入 prompt(RAG 投毒)
- 提出缓解措施:输入校验、检索过滤、工具沙箱、最小权限、护栏、策略提示、拒答逻辑、输出编码、监控告警
- 产出可审计、可用于决策的评测报告:方法、数据集、指标、阈值;定量与定性结果;风险评估摘要与建议的控制措施;向干系人清晰汇报残余风险、局限和 go/no-go 的理由
任职要求
- 数据科学、统计、工程、运筹等相关专业本科或硕士,学术背景优秀
- 4–7 年以上相关经验,覆盖以下一个或多个方向:ML/AI/GenAI/Agentic 工程(NLP/LLM)、评测工程、应用研究、安全测试/红队、设计保障安全可靠稳健的评测 harness
- Python 扎实,能构建评测 harness(数据处理、指标计算、编排、报告流水线)
- 理解 GenAI 系统架构:RAG、Embedding/向量搜索、prompt 编排、工具调用、多 Agent、记忆、路由
- 会设计指标和评测方法(rubric、自动打分、抽样策略、回归设计)
- 了解企业场景下 LLM 的风险与缓解(数据泄露、幻觉、prompt 注入、不安全内容、偏见)
- (强烈偏好)理解 OWASP Top 10 for LLM Applications 并能转成测试用例和控制措施;熟悉越狱、注入、工具误用、检索投毒等对抗测试方法;了解最小权限、安全工具调用、输出编码/校验等 secure-by-design 实践
- (偏好)熟悉评测框架和工具:RAGAS、DeepEval、LangSmith、Phoenix/Arize、自研 harness;A/B 思维、基线对比、样本量的统计严谨;可观测:结构化日志、OpenTelemetry、Langfuse 式 trace、仪表盘;Git、CI/CD、Docker、可复现环境
- 书面沟通强,能为技术和非技术读者写清评测方案与报告;敢于建设性质疑(「effective challenge」)并推动工程团队整改;能在快速变化的 GenAI 工具与风险环境中保持从容
- 另外要求:较强的口头、演示和主持能力;有项目管理经验;能同时协调多个项目
加分项 / 其他
- 鉴证/金融/监管环境经验(模型验证、风险接受流程、审计证据思维)
- 了解负责任 AI 框架(NIST AI RMF、ISO/IEC 42001、EU AI Act)
- 评测多语言系统或领域很深的企业助手
- Azure 生态(Azure OpenAI、AI Search、Function Apps、App Insights、Key Vault)
这份 JD 的看点:这是整组样本里对「评测工程」职责拆得最完整的一份,评测策略 → 数据集设计 → 指标 → LLM-as-judge 校准 → 自动化 harness → 红队 → CI 门禁 → 审计级报告,几乎可以直接当作评测工程师的能力地图。它把「Agent 行为」单独列成指标:工具调用准确率、步骤正确性、多余动作、循环检测,这是评测 Agent 和评测普通聊天机器人最大的区别。
✅ 你已经有的
- Python 生产级功底;评测 harness 本质是「批量跑用例 + 算指标 + 出报告」,你在稳定性测试平台里做过同类系统
- CI/CD 与发布门禁:「发布前回归门禁」对你是熟悉的工程动作
- ECharts 统计报告(覆盖率分布、甘特图):报告可视化你有基础
- 质量部背景 + AI 测试平台,对「怎么测非确定性系统」有第一手接触
⚠️ 差距
- RAGAS / DeepEval / LangSmith 这类评测框架的实操
- 统计严谨性(样本量、置信区间、一致性检验)
- OWASP LLM Top 10、红队与对抗测试
- 审计/合规文档的写法;JD 还要求较强的口头与演示能力
- 岗位在印度诺伊达,JD 未写远程
🎯 怎么补 / 怎么用
- 免费读完 OWASP Top 10 for LLM Applications,对你自己的 RAG 项目做一次「prompt 注入 + 工具误用」测试,写成一份报告——这是这类岗位最直接的作品
- 用 RAGAS 或 DeepEval 给你的 RAG 模块出 faithfulness / context precision 等指标
原帖:https://in.linkedin.com/jobs/view/agentic-ai-evaluation-engineer-at-ey-4432607497(岗位可能已下线)
查看英文原文(学英语对照用)
At EY, you’ll have the chance to build a career as unique as you are, with the global scale, support, inclusive culture and technology to become the best version of you. And we’re counting on your unique voice and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all. EY- Assurance – Senior – Digital Role : GenAI / Agentic AI Evaluation Engineer (Quality, Safety & Reliability) Position Details As part of EY GDS Assurance Digital, you will help design, build, and scale a standardized evaluation capability, focused on evaluating GenAI, RAG-based, and Agentic AI solutions before deployment. This role sits at the intersection of AI evaluation engineering, Responsible AI, and GenAI security/red teaming. The primary objective is to ensure GenAI/agentic systems are safe, reliable, robust, and fit-for-purpose, by designing evaluation strategies, building repeatable test harnesses, and generating auditable evidence that supports go/no-go decisions. You will work with global stakeholders (product teams, solution architects, risk & compliance, and assurance leadership) to define evaluation requirements, request test datasets from product teams, execute rigorous evaluations (functional + non-functional), and recommend mitigations and controls to reduce risk. This is a core full-time role that requires a hands-on AI Development mindset, strong evaluation mindset, and the ability to translate risk concerns into practical testing strategies and measurable acceptance criteria. Responsibilities Define and operationalize evaluation strategies for GenAI systems across use cases like Q&A assistants, summarization, extraction, drafting, agentic systems, and multi-step workflows. Translate business use-cases into a structured evaluation plan: scope, assumptions, success criteria, datasets, metrics, red-team scenarios, thresholds, and reporting requirements. Drive standardization: reusable evaluation templates, test case libraries, scoring rubrics, and reporting formats across product teams. Design structured dataset requirements for product teams and ensure coverage across: Core user journeys and primary business intents Edge cases (rare prompts, ambiguous queries, incomplete context) Adversarial cases (malicious prompts, jailbreak attempts, prompt injections) Bias & fairness cases (sensitive demographic proxies, protected attributes, stereotyping patterns) Define guidance for dataset sufficiency and statistical coverage (e.g., minimum samples, distribution balance, scenario matrices, stratification by intent/risk). Build reusable evaluation pipelines for: Answer quality (correctness, relevance, completeness, clarity) Grounding & faithfulness (RAG-specific: faithfulness, context precision/recall, hallucination rate, citation quality) Agentic behavior (tool-call accuracy, tool misuse, goal completion, step correctness, unnecessary actions, loop detection, safety of tool outputs) Operational quality (latency, cost/token budget, throughput, stability, retries, failure recovery) Combine LLM-as-judge and human evaluation in a calibrated way (rubric design, sampling plans, agreement checks). Implement automated evaluation harnesses in Python (preferred), enabling: batch runs on scenario suites configurable metric definitions reproducible runs with run IDs and artifacts storage of traces and outputs for auditability Execute structured red teaming aligned to OWASP Top 10 for LLM Applications, covering (examples): Prompt injection (direct + indirect) and tool hijacking Sensitive data disclosure / PII leakage Insecure output handling (downstream injection) Training data leakage / memorization probes Model denial-of-service / denial-of-wallet patterns Integrate evals into development lifecycle: pre-release regression gates, CI checks, benchmark comparisons across model versions/prompts/tools/retrievers. Perform adversarial testing for agentic workflows: tool misuse / over-permissioned tool access unauthorized action execution exfiltration via tools/connectors prompt injection via retrieved documents (RAG poisoning) Recommend mitigations: input validation, retrieval filtering, tool sandboxing, least-privilege permissions, guardrails, policy prompting, refusal logic, output encoding, monitoring alerts. Produce high-quality evaluation reports that are auditable and decision-ready, including: methodology, datasets, metrics, thresholds quantitative results qualitative results risk assessment summary and recommended control actions Present findings to stakeholders in a crisp, risk-informed manner; clearly explain residual risk, limitations, and rationale for go/no-go. Key Requirements/Skills & Qualification: Excellent academic background, including at a minimum a bachelor’s or a master’s degree in data science, Statistics, Engineering, Operational Research, or other related field with strong focus on modern data architectures, processes, and environments. 4–7+ years of relevant experience in one or more areas: ML/AI/GenAI/Agentic engineering (NLP/LLMs), evaluation engineering, applied research security testing / red teaming building and designing evaluation harness that ensures safety, reliability and robustness. Strong hands-on Python for building evaluation harnesses (data processing, metric computation, orchestration, reporting pipelines). Practical understanding of GenAI system architectures: RAG, embeddings/vector search, prompt orchestration, tool calling, multi-agent systems, memory, routing. Experience designing metrics and evaluation methods (rubrics, automated scoring, sampling strategy, regression design). Familiarity with LLM risks and mitigations, especially for enterprise contexts (data leakage, hallucinations, prompt injection, unsafe content, bias). Security / Red Teaming Skills (Strong Preference) Understanding of OWASP Top 10 for LLM Applications and how to translate it into test cases and controls. Experience with adversarial testing approaches: jailbreak prompts, injection patterns, tool misuse scenarios, retrieval poisoning patterns. Familiarity with secure-by-design practices for LLM apps: least privilege, safe tool invocation, output encoding/validation, monitoring. Evaluation frameworks and tooling: RAGAS, DeepEval, LangSmith, Phoenix/Arize, custom eval harnesses. Experimentation practices: A/B testing mindset, baseline comparisons, statistical rigor for sample sizes. Observability/tracing: structured logging, OpenTelemetry, Langfuse-style traces, dashboards. Basic DevOps practices: Git, CI/CD, containerization (Docker), reproducible environments. Strong written communication to produce clear evaluation plans and reports for technical + non-technical stakeholders. Ability to challenge assumptions constructively (“effective challenge”) and influence engineering teams toward remediation. Comfort operating in ambiguity with fast-evolving GenAI tooling and risk landscape. Preferred / Nice-to-Have Experience in Assurance/Finance/Regulatory environments (model validation, risk acceptance workflows, audit evidence mindset). Familiarity with responsible AI frameworks (NIST AI RMF, ISO/IEC 42001, EU AI Act concepts). Experience evaluating multilingual systems or domain-heavy enterprise assistants. Hands-on with Azure ecosystem (Azure OpenAI, AI Search, Function Apps, App Insights, Key Vault). Additional skills requirements: Excellent written, oral, presentation and facilitation skills Ability to coordinate multiple projects and initiatives simultaneously through effective prioritization, organization, flexibility, and self-discipline. Must have demonstrated project management experience. Knowledge of firm’s reporting tools and processes. Proactive, organized, and self-sufficient with ability to priorities and multitask. Analyses complex or unusual problems and can deliver insightful and pragmatic solutions. Ability to quickly and easily create/ gather/ analyze data from a variety of sources. A robust and resilient disposition able to encourage discipline in team behaviors What We Look For A Team of people with commercial acumen, technical experience, and enthusiasm to learn new things in this fast-moving environment An opportunity to be a part of market-leading, multi-disciplinary team of 7200 + professionals, in the only integrated global assurance business worldwide. Opportunities to work with EY GDS Assurance practices globally with leading businesses across a range of industries What Working At EY Offers At EY, we’re dedicated to helping our clients, from startups to Fortune 500 companies — and the work we do with them is as varied as they are. You get to work with inspiring and meaningful projects. Our focus is education and coaching alongside practical experience to ensure your personal development. We value our employees, and you will be able to control your own development with an individual progression plan. You will quickly grow into a responsible role with challenging and stimulating assignments. Moreover, you will be part of an interdisciplinary environment that emphasizes high quality and knowledge exchange. Plus, we offer: Support, coaching and feedback from some of the most engaging colleagues around Opportunities to develop new skills and progress your career The freedom and flexibility to handle your role in a way that’s right for you EY | Building a better working world EY exists to build a better working world, helping to create long-term value for clients, people and society and build trust in the capital markets. Enabled by data and technology, diverse EY teams in over 150 countries provide trust through assurance and help clients grow, transform and operate. Working across assurance, consulting, law, strategy, tax and transactions, EY teams ask better questions to find new answers for the complex issues facing our world today.
AI 评测需到岗海外★★★★★
Senior LLM Evaluation Engineer(AI Quality Engineer)
Aspire(IT 服务,约旦公司) · 埃及
黄金数据集回归评测幻觉检测质量评分卡验收标准
主要职责
- 为 LLM 应用设计并执行全面的评测策略
- 基于业务需求和经批准的参考数据,开发基准数据集、黄金数据集和评测场景
- 评估 AI 生成内容的:事实准确性、一致性、完整性、相关性、依据性(groundedness)、幻觉、无依据的断言、毒性与安全问题
- 进行人工评测和人在回路的 AI 回复评估
- 为生成式 AI 应用设计验收标准和质量指标
- 在模型、prompt、RAG 或配置变更后执行回归评测
- 识别并记录 AI 的失败、边界情况、推理错误、prompt 失败和检索问题
- 产出详细的评测报告、质量评分卡和发布建议
- 与 AI 工程师协作,根据评测结论改进 prompt、RAG 流水线、Embedding 和模型配置
- 对照可信知识源和业务规则验证 AI 回复
- 持续改进:构建可复用的评测数据集和测试资产
- 在各 AI 项目中定义 AI 质量标准、治理实践和评测方法论
任职要求
- 8 年以上软件质量、AI 质量工程、机器学习、数据科学或相关领域经验
- 评估过大语言模型或生成式 AI 应用
- 深入理解 LLM 行为、prompt 工程和 RAG
- 有识别幻觉、推理失败、事实错误和不一致输出的经验
- 设计过评测数据集、基准场景和验收标准
- 分析与批判性思维强,注重细节;有记录评测结果、质量指标和生产就绪评估的经验;沟通协作好
加分项 / 其他
- 熟悉 AI 评测框架:DeepEval、Ragas、LangSmith Evaluations、Promptfoo、OpenAI Evals 等
- 评测 RAG 系统与检索质量的经验;理解 LLM-as-a-Judge 与人工评测流程
- 在医疗、制药、法律或金融等强监管行业测试 AI 应用的经验
- 熟悉 Azure AI Foundry、Azure OpenAI、OpenAI API、Anthropic Claude、Gemini 等平台
- Playwright 或 API 测试工具经验;Python(偏好)
- AI 治理、负责任 AI、模型验证;向量库与知识检索;自动化评测流水线和 AI 可观测平台
这份 JD 的看点:它对传统测试/质量工程师最友好:不要求你会训练模型,核心是「定义好的标准、造好的数据集、发现问题、写发布建议」。把「质量评分卡」「发布建议」当成产出物,说明评测的价值是帮团队做上线决策。JD 里的技术关键词表(Technical Skills)也可以直接当作检索词。
✅ 你已经有的
- 质量部背景,你天然懂验收标准、回归测试、缺陷记录
- AI 测试平台里生成测试用例、管理测试资产,与这里「黄金数据集、测试资产」同类
- Python、API 测试能力
⚠️ 差距
- 评测框架(Promptfoo、Ragas、DeepEval)的实操
- 医药/法律等强监管领域经验
- 8 年「质量/AI 质量」的资历表述:你的年限够,但经历是开发,需要包装成质量视角
- 埃及地点,JD 未写远程
🎯 怎么补 / 怎么用
- 装 Promptfoo,给一个你熟悉的 LLM 应用写 20 条回归用例并跑起来,最快感受评测工作流
原帖:https://eg.linkedin.com/jobs/view/senior-llm-evaluation-engineer-at-aspire-jordan-4448029159(岗位可能已下线)
查看英文原文(学英语对照用)
Senior LLM Evaluation Engineer (AI Quality Engineer) About the Role We are seeking an experienced Senior LLM Evaluation Engineer to ensure the quality, accuracy, and reliability of Generative AI solutions used across our digital platforms. In this role, you will define and execute evaluation strategies for Large Language Models (LLMs), ensuring AI-generated content is factually correct, consistent, compliant, and production-ready. As the quality owner for AI-powered applications, you will design evaluation frameworks, create benchmark datasets, assess model performance, identify hallucinations and reasoning failures, and provide actionable feedback to improve prompts, retrieval pipelines, and model behavior. You will collaborate closely with AI Engineers, Product Managers, Medical, Legal & Regulatory (MLR), and Business stakeholders to ensure AI systems meet the highest quality standards before release. Key Responsibilities Design and execute comprehensive evaluation strategies for LLM-powered applications. Develop benchmark datasets, golden datasets, and evaluation scenarios based on business requirements and approved reference data. Evaluate AI-generated outputs for: Factual accuracy Consistency Completeness Relevance Groundedness Hallucinations Unsupported claims Toxicity and safety concerns Perform manual evaluations and human-in-the-loop assessments of AI responses. Design acceptance criteria and quality metrics for Generative AI applications. Execute regression evaluations following model, prompt, RAG, or configuration changes. Identify and document AI failures, edge cases, reasoning errors, prompt failures, and retrieval issues. Produce detailed evaluation reports, quality scorecards, and release recommendations. Collaborate with AI Engineers to improve prompts, RAG pipelines, embeddings, and model configurations based on evaluation findings. Validate AI responses against trusted knowledge sources and business rules. Support continuous improvement by building reusable evaluation datasets and test assets. Define AI quality standards, governance practices, and evaluation methodologies across AI initiatives. Required Qualifications 8+ years of experience in Software Quality, AI Quality Engineering, Machine Learning, Data Science, or related fields. Hands-on experience evaluating Large Language Models (LLMs) or Generative AI applications. Strong understanding of LLM behavior, prompt engineering, and Retrieval-Augmented Generation (RAG). Experience identifying hallucinations, reasoning failures, factual inaccuracies, and inconsistent AI outputs. Experience designing evaluation datasets, benchmark scenarios, and acceptance criteria. Strong analytical and critical thinking skills with exceptional attention to detail. Experience documenting evaluation results, quality metrics, and production readiness assessments. Excellent communication skills with the ability to collaborate across technical and business teams. Preferred Qualifications Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith Evaluations, Promptfoo, OpenAI Evals, or similar tools . Experience evaluating RAG systems and retrieval quality. Understanding of LLM-as-a-Judge methodologies and human evaluation workflows. Experience testing AI applications in regulated industries such as healthcare, pharmaceutical, legal, or finance. Familiarity with Azure AI Foundry, Azure OpenAI, OpenAI APIs, Anthropic Claude, Gemini, or similar platforms. Experience with Playwright or API testing tools is a plus. Technical Skills Large Language Models (LLMs) Generative AI LLM Evaluation AI Quality Assurance Prompt Engineering Prompt Evaluation Retrieval-Augmented Generation (RAG) Hallucination Detection Groundedness Evaluation Benchmark Dataset Creation Human-in-the-Loop Evaluation Regression Evaluation AI Safety & Compliance API Testing Python (preferred) Azure AI Foundry (preferred) LangSmith, Ragas, DeepEval, Promptfoo, or equivalent evaluation frameworks Nice to Have Experience with healthcare or regulated AI solutions. Knowledge of AI governance, responsible AI, and model validation practices. Experience working with vector databases and knowledge retrieval systems. Familiarity with automated evaluation pipelines and AI observability platforms. Experience collaborating with cross-functional AI engineering and product teams. What You'll Bring You are passionate about ensuring AI systems are trustworthy, accurate, and safe. You understand that building an LLM is only part of the solution—evaluating, validating, and continuously improving its outputs is equally important. You enjoy combining analytical thinking with human judgment to ensure AI applications deliver reliable, production-ready experiences.
AI 评测需到岗海外★★★★★
AI Evaluation Engineer
Dialpad · 加拿大 · 安大略省基奇纳
💰 安大略省基础薪资 CAD $96,000–$116,250LLM-judge 校准回归评测A/B 对比红队分析数据标注
主要职责
- 为 Agent、NLP 和语音工作流设计并执行验证策略,覆盖 staging、beta 和候选发布版本
- 构建、运行并改进回归评测、A/B 对比和红队分析,判断产品和模型变更是否可以推进
- 共同负责 LLM-judge 指标的开发、校准和 prompt 打磨(跨多个评测维度)
- 创建、配置并监控数据标注任务,保证评测集和校准数据集按计划持续补充
- 开发并维护 QA 工具、notebook 和流水线组件,让周期性评测可规模化、可跨团队复用
- 调查 bug、分诊问题,决定是升级给工程、补进测试集,还是做后续分析
- 与应用科学、工程和产品 QA 等跨职能团队协作
任职要求
- 计算机、软件工程、计算语言学等相关专业本科或硕士
- 3 年以上 QA、测试工程、模型评测或 AI 驱动产品的应用 ML 质量经验
- 有跨手工和自动化流程设计结构化测试策略的经验
- 能在语音、NLP、LLM 或 Agent 等复杂 AI 系统上工作
- 有评测数据集、gold set、对抗测试集或 AI 基准创建的经验
- 分析能力强,善于调查失败、对比输出、找出可行动的质量模式
- 能与跨职能技术团队协作,通过文档和报告清晰沟通
这份 JD 的看点:重点是「LLM-judge 指标的开发与校准」:用大模型给输出打分不是拿来就用,要和人工标注对照、校准,才能作为发布依据。另外「错误分析后决定升级、补测试集、还是继续分析」这一句,把评测工程师的日常工作流讲清楚了。要求只有 3 年经验,是这组里门槛最低的评测岗。
✅ 你已经有的
- 质量部背景,QA/测试工程正是你的本行
- 测试资产管理系统(知识库、用例库)服务 50+ 团队,和「评测数据集维护」类似
- AI 测试平台里做过用例生成,接触过 AI 输出的质量问题
⚠️ 差距
- LLM-judge 的校准方法(和人工标注的一致性)
- 语音/NLP 评测经验
- 数据标注任务的运营
- 加拿大到岗,JD 未写远程
🎯 怎么补 / 怎么用
- 做一个小实验:拿 50 条 LLM 回答,自己先人工打分,再让另一个 LLM 打分,算一致率,写下不一致的原因——这就是「校准」的最小版本
原帖:https://ca.linkedin.com/jobs/view/ai-evaluation-engineer-at-dialpad-4444461503(岗位可能已下线)
查看英文原文(学英语对照用)
About Dialpad Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage. Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved. Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust Dialpad. Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile. Being a Dialer At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more. We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves. We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . Your role As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside our existing evaluation lead. A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions. This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office. What You’ll Do You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates. You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward. You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions. You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule. You will develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams. You will investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis. You will collaborate with cross-functional teams, including applied science, engineering, and Product QA. Skills You’ll Bring Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field. 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products. Experience designing structured test strategies across manual and automated workflows. Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products. Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems. Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns. Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting. For exceptional talent based in Ontario, Canada the target base salary range for this position is posted below. Our salary ranges are determined by role, level, and location. The range displayed on each job posting reflects the target range for new hire salaries for the position. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your preferred location during the hiring process. Please note that the compensation details listed in Ontario role postings reflect the base salary only, and do not include bonus, equity, or benefits. Ontario Salary Range $96,000—$116,250 CAD Why Join Dialpad Work at the center of the AI transformation in business communications Build and ship agentic AI products that are redefining how companies operate Join a team where AI amplifies every employee’s impact Competitive salary, comprehensive benefits, and real opportunities for growth We believe in investing in our people. Dialpad offers competitive benefits and perks, cutting-edge AI tools, and a robust training program that help you reach your full potential. We have designed our offices to be inclusive, offering a vibrant environment to cultivate collaboration and connection. Our exceptional culture, repeatedly recognized as a Great Place to Work, ensures that every employee feels valued and empowered to contribute to our collective success. Don’t meet every single requirement? If you’re excited about this role and possess the fundamental traits, drive, and strong ambition we seek, but your experience doesn’t meet every qualification, we encourage you to apply. Dialpad is an equal-opportunity employer. We are dedicated to creating a community of inclusion and an environment free from discrimination or harassment.
AI 评测需到岗海外★★★★★
Member of Technical Staff, North Modelling (Evals)
Cohere · 英国 · 伦敦
评测策略Agent 评测反馈闭环eval 数据来源应用型 MLE
主要职责
- 负责 North 的评测策略:定义需要度量什么,覆盖 Agent 工作流、工具使用、企业知识工作(深度研究、文档创建/编辑)和其他人机协作
- 从产品现实出发构建高质量 eval:用户反馈、生产故障、保护隐私的使用日志、内部试用(dogfooding)、客户需求和产品目标
- 建立系统,持续把 North 从用户、客户和功能团队那里学到的东西转成 eval,让度量跟上产品,而不是变成静态基准
- 作为「North 在建模团队里的代言人」:从评测结果、客户数据和产品背景里提炼洞察,把模型失败、能力缺口和 North 的特殊需求转成给中央建模团队的可执行建议
- 用这些 eval 指导模型选型、补丁和定期的模型更新
任职要求
- (适合你如果)通过 eval、反馈闭环、数据整理、prompt、模型适配或模型选型,改进过基于 LLM / Agent / AI 产品的系统
- 把评测当作一门手艺:有代表性的任务、精确的 rubric、干净的数据、失败分析、回归追踪,并且知道什么时候某个指标在制造「虚假信心」
- 有很强的应用型 MLE 判断力,能清楚地推理模型行为、评测有效性、产品结果和生产取舍
- 能贴近用户和产品团队工作,把杂乱的定性信号转成其他建模团队能用的度量
- 自驱、务实,喜欢需要技术深度和产品理解并存的开放问题
- 在乎让 AI 在真实企业场景里真正有用,而不只是在抽象基准上更好
这份 JD 的看点:最有启发的一句:「评测要让度量跟上产品,而不是变成静态基准」。也就是把用户反馈、线上失败、使用日志持续转成新的评测用例。另一句同样重要:知道「什么时候某个指标在给你虚假信心」,这是评测工作最难的判断力。
✅ 你已经有的
- Agent、工具调用、MCP 的产品化经验
- 从线上问题回流到测试集的质量工作思路
⚠️ 差距
- 应用型 MLE 的建模判断(模型适配、选型、微调)
- 评测有效性的统计与方法论
- 英国伦敦、跨欧美时区协作
🎯 怎么补 / 怎么用
- 把这个思路用在自己的项目:每次线上发现一个 Agent 失败案例,就补进回归集,并记录「这个案例为什么之前没被发现」
原帖:https://uk.linkedin.com/jobs/view/member-of-technical-staff-north-modelling-evals-at-cohere-4454992608(岗位可能已下线)
查看英文原文(学英语对照用)
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why This Role? North is Cohere’s AI workspace platform for enterprises: a secure, customizable environment where companies can use AI across their real workflows while maintaining control over sensitive data. North connects AI agents with workplace tools, applications, and business context, helping users delegate complex work, build automations, inspect outputs, and collaborate with AI in production environments. As North becomes more capable, one of the most important questions is also one of the hardest: how do we know whether the model is actually getting better for the workflows customers care about? This role is about being the voice of North inside modelling. You will build the evaluation systems, feedback loops, and applied modelling workflows that make sure model progress translates into better product outcomes for North users. You will work closely with North product teams, customer-facing teams, and modelling teams to define what “good” means across the product surface, turn real usage and product direction into high-quality evals, and use those evals to guide model selection, patches, and regular model updates. This is neither a pure research role nor a conventional product engineering role. It is a rigorous applied MLE role for someone who cares deeply about measurement, model behavior, and real-world product quality. You should be excited by the craft of building careful evals: evals that capture messy agentic workflows, reflect actual customer needs, resist superficial benchmark hacking, and provide useful signal for where the product and models need to go next. As a Member of Technical Staff, North Modelling (Evals), You Will: Own the eval strategy for North: define what we need to measure across agent workflows, tool use, enterprise knowledge work such as deep research, document creation or editing, and other human-AI interactions. Build high-quality evals from the realities of the product: user feedback, production failures, privacy-preserving usage logs, internal dogfooding, customer needs, and forward-looking product goals. Create systems that continuously turn what North is learning from users, customers, and feature teams into evals, so measurement keeps pace with the product rather than becoming a static benchmark. Be the voice of North inside modelling: extract clear insight from eval results, customer-derived data and product context, then turn model failures, capability gaps, and North-specific needs into actionable recommendations for the central modelling teams. You May Be a Good Fit If You have improved LLM-powered, agent-powered, or AI-product systems through evals, feedback loops, data curation, prompting, model adaptation, or model selection. You care deeply about evaluation as a craft: representative tasks, precise rubrics, clean data, failure analysis, regression tracking, and knowing when a metric is giving false confidence. You have strong applied MLE judgment and can reason clearly about model behavior, eval validity, product outcomes, and production tradeoffs. You are comfortable working close to users and product teams, while translating messy qualitative signals into measurement that other modelling teams can act on. You are self-directed, practical, and motivated by open-ended problems where the right answer requires both technical depth and product understanding. You care about making AI systems genuinely useful in real enterprise settings, not just better on abstract benchmarks. If some of the above does not line up perfectly with your experience, we still encourage you to apply. Location This team works closely across Europe and East Coast North America time zones. We are open to candidates who can collaborate effectively within those hours. Full-Time Employees At Cohere Enjoy These Perks A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch. Full health and dental benefits, including a separate budget for mental health. RRSP matching, 401K, Pension Scheme. 100% Parental Leave top-up for up to 6 months, for either parent. Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace improvement credit. Education & learning stipend for conferences, courses, and coaching. 6 weeks of paid vacation (30 working days!) Budget for traveling to other offices if you are remote, plus an annual company offsite. How And Where We Work Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon. For those in the office: a daily lunch program, plenty of snacks, and regular community and social events. For those not near an office: a co-working benefit so you can work alongside others in your city. Everyone receives a $500 home office stipend to set up your workspace properly. If any of the above doesn’t line up exactly with your experience, we still encourage you to apply. We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs. We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider. Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers page.
AI 评测需到岗海外★★★★★
Senior Research Scientist, Model Evaluation
Cohere · 英国 · 伦敦
评测基准训练 LLM judge数据合成评测方法研究
主要职责
- 创建有雄心的新评测基准,推到模型能力的边界
- 在高度跨职能的团队里,把模型反馈转成可信、可重复的评测
- 研究推进 LLM 评测方法的前沿:训练 LLM judge、改进基于 LLM 的数据合成流水线、提升评测效率
- 构建可扩展、可复用的工具,深入分析模型表现
任职要求
- 喜欢快速做原型,展示 LLM 能力的边界,并且开发过度量这些能力的资源
- 花过几十个小时审阅复杂数据和 LLM 输出,保证数据质量
- 对严谨度量 AI 能力很执着,也执着于「你的度量是否真的对应你在乎的能力」
- 有很强的软件工程能力
这份 JD 的看点:它开头一句很关键:「当模型在许多真实场景里已经超越人类,我们必须发明新的评测技术,才能准确反映模型已经能做什么」。评测本身已经成为一个研究问题。要求里「花几十个小时审阅数据」,说明看数据是评测工作的基本功。
✅ 你已经有的
- 强软件工程能力是共同项
⚠️ 差距
- 研究背景:评测方法论、LLM judge 训练、合成数据
- 论文与基准构建的经历
- 伦敦到岗
🎯 怎么补 / 怎么用
- 作为方向了解即可;能做的低成本练习是「看数据」:定期抽 100 条模型输出,自己分类失败原因,练判断力
原帖:https://uk.linkedin.com/jobs/view/senior-research-scientist-model-evaluation-at-cohere-4317721089(岗位可能已下线)
查看英文原文(学英语对照用)
Who are we? Cohere is the leading security-first enterprise AI company. We build cutting-edge foundation AI models and end-to-end products that are designed to solve real-world business problems. We’re training and deploying frontier models for enterprises who are building AI systems. We believe that our work is instrumental to the widespread adoption of AI and we are looking for folks that want to be part of that. We obsess over what we build. Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers. Cohere is a team of researchers, engineers, designers, and more, who are all passionate about their craft. We are a global technology company headquartered in Toronto with key offices in London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul. Join us! Why this role? Evaluation is critical to making progress in scaling intelligence. As models continue to become superhuman in many real-world use cases, we must continue to develop new evaluation techniques that accurately reflect what models are already capable of, as well as set the agenda for what future models should be capable of. In this role, you are responsible for creating these next-generation evaluation methods and infrastructure to measure LLM progress. As a Senior Research Scientist, Model Evaluation, You Will Create ambitious new evaluation benchmarks that push the limits of what our models can accomplish. Work on highly cross-functional teams to translate model feedback into trustworthy, repeatable evaluations. Conduct research to advance the state-of-the-art in LLM evaluation methods, including training LLM judges; refining LLM-based data synthesis pipelines; and improving evaluation efficiency. Build scalable and reusable tools for digging into model performance. You May Be a Good Fit If You enjoy rapidly building prototypes that demonstrate the boundaries of what LLMs are capable of, and you have developed resources to measure those capabilities. You have spent dozens of hours reviewing complex data and LLM outputs to ensure high data quality. You are obsessive about rigorously measuring AI capabilities, and also about making sure your measurements actually align with the capabilities you care about. You have strong software engineering skills. Full-Time Employees At Cohere Enjoy These Perks A weekly lunch stipend of $75/£75 or equivalent in your local currency for lunch. Full health and dental benefits, including a separate budget for mental health. RRSP matching, 401K, Pension Scheme. 100% Parental Leave top-up for up to 6 months, for either parent. Annual enrichment benefits: Arts & culture, fitness/wellness, quality time, and a workspace improvement credit. Education & learning stipend for conferences, courses, and coaching. 6 weeks of paid vacation (30 working days!) Budget for traveling to other offices if you are remote, plus an annual company offsite. How And Where We Work Cohere is remote-friendly, but we also have offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin and Seoul with more opening soon. For those in the office: a daily lunch program, plenty of snacks, and regular community and social events. For those not near an office: a co-working benefit so you can work alongside others in your city. Everyone receives a $500 home office stipend to set up your workspace properly. If any of the above doesn’t line up exactly with your experience, we still encourage you to apply. We strive to create an inclusive work environment for all; we welcome applicants from all backgrounds and are committed to providing equal opportunities. Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form, and we will work together to meet your needs. We may use AI-enabled tools to screen and assess applicants against the criteria for this position. This helps our recruiters identify potentially qualified candidates, but it doesn't limit the applications our recruiters may review or consider. Beware of Scams: Cohere will never ask for payment or third-party services (e.g., CV writing) as part of our hiring process. All legitimate roles are listed on the Cohere careers page and LinkedIn only, with all communications from Cohere employees coming from an @cohere.com or @cw.cohere email alias. If jobs are viewed on other sites then please verify these through our official careers page.
AI 评测限定国家★★★★★
AI Quality Engineer
Dealstitch AI · 美国
步骤级评测CI/CD 门禁LLM-as-judgeDeepEval/RAGAS黄金数据集
主要职责
- 负责为每一次变更把关的评测流水线,保持自动化并跑在 CI/CD 里,而不是靠手工抽查
- 按「步骤级」给 Agent 打分,而不只看最终答案:工具选择、规划与推理链、检索质量,让细微的回归在客户发现之前浮出水面
- 随着新失败模式出现扩展评测工具集:LLM-as-a-judge、RAG 忠实度与相关性、幻觉检测、prompt 与 Agent 回归测试集、面向安全的红队
- 维护镜像真实生产输入的黄金数据集,以及 Agent 上线后的追踪与可观测
- 负责发布质量线,并与工程团队一起持续提高它
任职要求
- 扎实的软件工程,Python;懂得测试非确定性系统——传统的断言式测试不够用
- 评测过 LLM 和 Agent 系统:Agent 图(如 LangGraph)、评测框架(DeepEval、RAGAS、G-Eval、LangSmith 或同类)、LLM-as-a-judge 方法
- 熟悉 RAG 与检索质量,以及生产环境中 Agent 的可观测与追踪
- 重视度量:先定义什么叫「好」,再搭出证明它的 harness
加分项 / 其他
- 数据流水线测试经验,或接触过私募市场、金融服务、企业数据
这份 JD 的看点:这份 JD 值得反复读:「按步骤打分(选工具、规划、检索),而不是只看最终答案」——这是 Agent 评测和聊天评测最大的不同。「传统的断言式测试不够用」——一句话点出了非确定性系统测试的难点。整个 JD 只有几段,是入门评测最好的一张地图。
✅ 你已经有的
- Python 生产级;CI/CD 与发布门禁是你的主场(50+ 业务线接入)
- 质量部背景 + AI 测试平台,「测试」对你不是新概念,只是对象从确定性系统换成 LLM/Agent
- 对「回归测试集镜像真实生产输入」这类做法有直觉
⚠️ 差距
- DeepEval / RAGAS / G-Eval / LangSmith 的实战使用
- LangGraph 的 Agent 图与步骤级追踪
- LLM-as-a-judge 的校准
- 美国地点,是否接受境外要问
🎯 怎么补 / 怎么用
- 两周计划:① 用 LangGraph 搭一个 3 步的小 Agent;② 用 DeepEval 对最终答案和「工具选择」分别打分;③ 接进 GitHub Actions,指标下降就红灯;④ 写成 README。做完就是一份对得上这个岗位的作品
原帖:https://www.linkedin.com/jobs/view/4472485386/(岗位可能已下线)
查看英文原文(学英语对照用)
About The Role An agent that demos well is not the same as an agent that works in production, so every agent we ship is backed by evals. We're hiring an AI Quality Engineer to own that system. You'll run and extend the evaluation pipelines that prove our agents work, both in client engagements and across our own internal agent fleet, and keep them honest as the agents, models, and data underneath them change. What you'll do Own the evaluation pipelines that gate every change we ship, keeping them automated and running in CI/CD rather than leaning on manual spot-checks. Keep agents scored at the step level, not just the final answer, across tool selection, planning and reasoning chains, and retrieval quality, so subtle regressions surface before clients do. Extend the eval toolkit as new failure modes appear: LLM-as-a-judge, RAG faithfulness and relevance, hallucination detection, prompt and agent regression suites, and red-teaming for safety. Maintain the golden datasets that mirror real production inputs and the tracing and observability that watch our agents once they are live. Own the quality bar for what ships and raise it over time with the rest of the engineering team. What we're looking for Strong software engineering with Python, and a feel for testing non-deterministic systems, where classic assertion-based tests are not enough. Experience evaluating LLM and agent systems: agent graphs (e.g. LangGraph), eval frameworks (DeepEval, RAGAS, G-Eval, LangSmith, or similar), and LLM-as-a-judge methods. Fluency with RAG and retrieval quality, plus observability and tracing for agents in production. A bias for measurement: you define what "good" means, then build the harness that proves it. Bonus: data-pipeline testing experience, or exposure to private markets, financial services, or enterprise data. Why Dealstitch We're a senior team with 40+ years of combined private-market technology experience. You'll work on problems that matter, with direct access to decision-makers, and see the impact of your work in weeks, not quarters.
AI 评测地区开放★★★★★
Open Source AI Engineer (TypeScript)
Arize AI · 全球远程
💰 $185,000–$200,000LLM 评测可观测TypeScript开源Agent harness
主要职责
- 构建 LLM 可观测框架:设计、架构并开源新的库、流水线和 API,让监控、评估、改进 LLM 输出更简单
- 与社区协作:与开源生态紧密合作,收集反馈、评审 PR、根据真实开发者需求把握项目方向
- 快速原型与迭代:试验最新的 LLM 技术,把研究变成实用的开发者工具
- 改进可观测与调试:与现有平台集成,展示对 LLM 行为更深入的洞察,帮助团队快速诊断和修复幻觉、偏见等问题
- 布道与教育:写博客、白皮书、教程和文档,帮助开发者用好开源工具,壮大 LLM 评测社区
任职要求
- 开源拥护者:相信协作与社区驱动能带来最好的创新
- 有创造性的问题解决能力:喜欢处理模糊问题,找优雅的技术方案
- 重数据与指标:重视实证结果,喜欢创造或改进评测指标,并根据真实反馈迭代
- 技术好奇心强,不断学习新的 LLM 架构、prompt 策略或新的库标准
- 构建者心态:享受把想法从原型做到用户喜欢的生产级方案
- 动手做过 LLM:熟悉主流 LLM 框架、prompt 技巧和 Agent harness
- TypeScript 精通:能写在客户端、服务端和各种边缘运行时无缝运行的同构代码
- 开源履历:有开源项目贡献、带有趣 AI 演示的个人 GitHub、或活跃的开发者社区参与
- (加分)ML 可观测与工具:熟悉调试 AI 应用、探索 Embedding、构建数据密集型看板
这份 JD 的看点:要看这份 JD 是因为:它把「评测」当作产品和开源项目来做,而不是内部质量流程。要求里对「开源履历」和「布道」的强调也很明确,说明这类岗位的招聘方式是「看你公开做过什么」。薪资和远程范围都是这组里最透明的。
✅ 你已经有的
- TypeScript / React 18 是你的强项(★5),前端复杂交互(React Flow 1000+ 节点、ECharts)经验很扎实
- 做过 LLM 应用、MCP、Agent 框架(nanobot),对 harness 有认识
- 数据可视化 + 可观测(你做过覆盖率分布图、任务甘特图、环境拓扑图),和「构建数据密集型看板」直接相关
⚠️ 差距
- 没有公开的 GitHub 项目——这是这类岗位最大的门槛,也是你自己计划要做的事(开源 nanobot、flower-pc-fontend)
- 同构 TS 代码(Node + 浏览器 + 边缘运行时)
- 评测指标的设计与开源社区运营经验
🎯 怎么补 / 怎么用
- 优先级最高:把 nanobot 开源,配上评测/追踪功能(哪怕只是接 OpenTelemetry 记录每步 trace),写 README 和一篇博客。这份 JD 招的就是这种人
原帖:https://wellfound.com/jobs/4523736-open-source-ai-engineer-typescript(岗位可能已下线)
查看英文原文(学英语对照用)
The Opportunity AI is rapidly transforming the world. Whether it’s developing the next generation of human-level intelligence, enhancing voice assistants, or enabling researchers to analyze genetic markers at scale, AI is increasingly integrated into various aspects of our daily lives. Arize AI is the leading AI observability and evaluation platform, empowering AI engineers to build and deploy high-performing, reliable models. As the AI landscape shifts from traditional ML to generative AI and agentic systems, Arize ensures teams have the tools to monitor, troubleshoot, and improve AI in production. We’re looking for an Open Source AI Engineer to join our growing OSS team to drive the development of new frameworks, metrics, and tooling that help people build, test, and improve LLM tasks. You’ll play a lead role in shaping how developers measure and understand performance in advanced AI systems, all in the open. What You’ll Work On Build LLM Observability Frameworks: Design, architect, and open-source new libraries, pipelines, and APIs that make it simpler to monitor, evaluate, and improve LLM output. Collaborate with the Community: Partner closely with the broader AI open source ecosystem, gather feedback, review pull requests, and steer the direction of the project to address real developer needs. Prototype and Iterate Rapidly: Experiment with state-of-the-art LLM techniques, turning research into practical developer tooling. Improve Observability and Debugging: Integrate with our existing platform to surface deeper insights on LLM behavior—help teams quickly diagnose and fix issues such as hallucinations or bias. Educate and Evangelize: Write blog posts, white papers, tutorials, and documentation to help developers succeed with our open source tools and grow the LLM eval community. What We’re Looking For We’re looking for an engineer who’s deeply passionate about AI, loves working in the open, and thrives in a fast-paced environment where “everyone wears multiple hats”. You likely share our core values: Open Source Champion: You believe collaboration and community-driven development unlocks the best innovations. Creative Problem Solver: You enjoy tackling ambiguous challenges and finding elegant technical solutions. Data & Metrics Driven: You value empirical results, enjoy creating or refining evaluation metrics, and iterate based on real-world feedback. Technically Curious: You’re always learning—exploring new LLM architectures, prompt engineering strategies, or emerging library standards. Builder Mindset: You relish the process of taking ideas from initial prototypes to production-ready solutions that delight users. Desired Skills & Experience Hands-on LLM Experience: Familiarity with popular LLM frameworks, prompt engineering techniques, and agent harnesses. Strong TypeScript Proficiency: You understand TypeScript and can write isomorphic code that executes seamlessly across clients, servers, and diverse edge runtimes. Open Source Track Record: Contributions to open source projects, personal GitHub repos with interesting AI demos, or a history of active engagement in developer communities. ML Observability & Tools: Familiarity with debugging AI applications, exploring embeddings, or building data-heavy dashboards is a plus. Why Work With Us Shape the Future of AI Evaluation: Be at the forefront of designing new ways to measure and improve next-generation LLMs. High Impact, Real Ownership: Join a team that values autonomy and speed. You’ll drive major initiatives from day one and see your work used by developers worldwide. Fully Remote, Flexible Environment: We are a fully remote company with offices in the Bay Area and NYC for those who prefer in-person collaboration. Cutting-Edge Challenges: Our platform already helps analyze millions of AI predictions daily, giving you the chance to refine your evaluation tooling on real, large-scale production workloads. Work With a Talented, Passionate Team: Collaborate closely with top engineers who are dedicated to making AI more transparent, reliable, and impactful. The estimated annual salary and variable compensation for this role is between $185,000 - $200,000, plus a competitive equity package. Actual compensation is determined based upon a variety of job related factors that may include: transferable work experience, skill sets, and qualifications. Total compensation also includes a comprehensive benefit package, including: medical, dental, vision, 401(k) plan, unlimited paid time off, generous parental leave plan, and others for mental and wellness support. While we are a remote-first company, we have opened offices in New York City and the San Francisco Bay Area, as an option for those in those cities who wish to work in-person. For all other employees, there is a WFH monthly stipend to pay for co-working spaces.
AI 评测地点待确认★★★★★
Research Engineer, Coding Evaluation & Training Data
Surge AI · 地点未写
编码评测RL 环境rubric/golden set奖励信号沙箱/容器
主要职责
- 端到端负责编码数据项目,从最初的范围界定、试点设计,到执行、迭代和规模化
- 设计 Agent 化的训练工作流和任务结构,镜像真实的软件工程工作(重构、调试、代码评审、大型代码库导航、工具使用)
- 定义并迭代 rubric、golden set 和奖励信号,捕捉真正的工程价值,而不只是「能不能编译通过」
- 以很强的软件工程品味评估数据和标注者的产出,判断哪些达到了前沿训练的标准
- 为编码类标注者设计并执行资格认证流程,包括对编码能力的实操评估
- 搭建或配合搭建复杂的技术环境(容器、代码仓库、测试 harness、沙箱、代码执行基础设施)
- 与客户方的技术人员合作,把高层次的训练目标转成具体的项目和技术环境
- 与 Surge 的工程、产品、运营紧密合作,改进编码数据产品、内部工具和执行流程
任职要求
- 3–6 年以上专业软件工程经验,构建并维护过真实系统
- 至少精通一种主流语言,习惯在生产代码库里工作
- 对好的工程有很高的「品味」:在乎正确性、代码质量,以及真实的工程团队是怎么工作的
- 能推理并调试技术环境,包括容器、依赖和自动化测试配置
- 想把项目端到端负责到底,包括范围界定、流程设计、执行和持续改进
- 书面和口头沟通能力出色,能与客户的资深工程师建立可信的对话,把模糊的目标变成可执行的计划
- 对 AI/ML 系统充满兴趣,认同数据、评测和奖励设计对提升 Agent 编码能力的作用
这份 JD 的看点:最值得读的一句是「奖励信号要捕捉真正的工程价值,而不只是能不能编译」。这是评测设计的核心难点:怎么用可自动判定的标准去逼近「好代码」这件很难量化的事。JD 把 rubric、golden set、reward signal 放在同一层,说明评测与训练信号在前沿实验室里是同一回事。
✅ 你已经有的
- 9 年软件工程,构建并维护过大型真实系统(腾讯云 50+ 业务线、千级并发)
- 容器/测试 harness/沙箱环境:环境管理平台、稳定性测试平台,正好是「搭复杂技术环境」这类工作
- 质量部背景,对「怎么判定一份产出合格」有职业直觉
⚠️ 差距
- 与客户资深工程师对话的口头沟通(这是本岗位的明确要求)
- 奖励设计与 RL 环境的理论背景
- 标注者资格认证/数据运营的经验
- 地点与薪资未知
🎯 怎么补 / 怎么用
- 选一个你熟悉的任务(如「修一个小 bug」),设计 5 条 rubric,再故意写几份「能通过测试但写得很烂」的答案,看你的 rubric 能不能区分——这就是这个岗位每天的思考
原帖:https://surgehq.ai/careers/research-engineer-coding-evaluation-training-data(岗位可能已下线)
查看英文原文(学英语对照用)
The Role As a Research Engineer, Coding Evaluation & Training Data at Surge, you'll sit at the intersection of software engineering and product to build and run the systems that teach frontier models how to code. You’ll own end-to-end coding data projects for top AI labs, designing tasks, RL environments, and evaluation schemes that reflect real-world software engineering. This is an ideal role for a software engineer who wants to keep using their engineering skill set every day, but is most excited by training data quality, agentic evaluation, and system design (humans + models + tools) rather than traditional product feature work. What You'll Do Own end-to-end coding data projects, from initial scoping and pilot design through execution, iteration, and scale-up. Design agentic training workflows and task structures that mirror real-world SWE work (e.g., refactoring, debugging, code review, large-repo navigation, tool use). Define and iterate on rubrics, golden sets, and reward signals that capture true engineering value, not just “does it compile.” Evaluate data and worker output with strong SWE taste, and make calls about what meets the bar for frontier training. Design and run qualification processes for coding workers, including hands-on assessments of their coding ability. Set up or partner on complex technical environments (e.g., containers, repos, test harnesses, sandboxes, code execution infrastructure). Partner with technical staff at our clients to translate high-level training goals into concrete projects and technical environments. Collaborate closely with Surge engineering, product, and operations to improve our coding data products, internal tools, and execution processes. What We're Looking For 3–6+ years of professional software engineering experience building and maintaining real systems. Strong coding ability in at least one mainstream language and comfort working in production codebases. High “taste” for good engineering: you care about correctness, code quality, and how real engineering teams actually work. Ability to reason about and debug technical environments, including containers, dependencies, and automated test setups. Interest in owning projects end-to-end, including scoping, workflow design, execution, and continuous improvement. Excellent written and verbal communication skills, with the ability to speak credibly with senior client engineers and translate fuzzy goals into actionable plans. Excitement about AI/ML systems and the role of data, evaluation, and reward design in improving agentic coding capabilities. How to Apply We created a short exercise to help us get to know you better. Once you’ve completed it, please email your submission and resume to careers@surgehq.ai and include the role you’re applying for in the subject line.
AI 评测地点待确认★★★★★
Backend Engineer(RL 环境 / 训练基础设施)
Surge AI · 地点未写
容器化环境高并发 Agent训练基础设施数据流水线可观测
主要职责
- (JD 没有单列职责,用两个「我们在解决的问题」说明工作内容)
- 问题一:怎样架构大规模模拟环境,托管数百个并发 Agent——每个 Agent 都生活在自己的容器化世界里——同时不牺牲稳定性和速度?
- 问题二:怎样构建基础设施,支撑高吞吐的模型训练循环并扩展到数千个环境,同时保证每个 Agent 的世界一致、高性能、并且能实时观测?
任职要求
- 这是资深、动手型的工程岗位,适合想要完整所有权的人
- 热爱构建健壮、可扩展的后端系统
- 有设计并实现高效数据处理流水线、API 和数据库架构的经验,能处理大规模数据
- 对性能优化有热情,享受解决复杂后端挑战、把系统优化到又快又高效
- 在乎我们要解决的问题:让下一代 AI 模型更安全、更有用
- 期待加入小而成长的团队,重视自主性,希望大家做出职业生涯最好的工作
这份 JD 的看点:JD 只写了两个具体问题,但很有信息量:「数百个 Agent,各自一个容器」「数千个环境」「每个世界一致且可观测」。这其实是分布式调度 + 容器编排 + 状态一致性问题,和 Agent 应用开发是两个层次的工程。这个岗位是「评测/训练的地基」,让你不写模型也能站在前沿实验室的核心。
✅ 你已经有的
- Go 后端、K8s、Docker 是你的主力;环境管理平台就是「统一调度多套 K8s 环境、动态创建/回收」,和「数千个容器化环境」是同类问题
- Temporal 分布式工作流:任务状态持久化、故障恢复、千级并发,可直接迁移到「环境编排」
- 数据处理流水线、gRPC/RESTful API、监控运维
⚠️ 差距
- JD 极简,具体技术栈和年限要进一步问
- 训练循环、RL 环境的领域知识(需要补背景)
- 地点与薪资未知
🎯 怎么补 / 怎么用
- 做一个小实验:用 Docker + Temporal 起 50 个隔离容器,每个跑一个简单 Agent 环境,写一页文档讲清「怎么保证环境一致、怎么观测、怎么回收」。作品可以直接放进申请邮件
原帖:https://surgehq.ai/careers/backend-engineer(岗位可能已下线)
查看英文原文(学英语对照用)
What We’re Looking for This is a senior, hands-on engineering role for builders who want full ownership. You love building robust, scalable backend systems. You're experienced in designing and implementing efficient data processing pipelines, APIs, and database architectures that can handle large volumes of data. You're passionate about performance optimization. You enjoy tackling complex backend challenges and optimizing systems for speed and efficiency. You care about the problems we’re trying to solve . You're interested in helping make the next generation of AI models safer and more useful. You’re excited about joining a small, growing team. We value autonomy and empower people to do the best work of their career. Some Examples of the Problems We’re Solving What’s the best way to architect large-scale simulated environments that can host hundreds of concurrent agents—each living in its own containerized world—without sacrificing stability or speed? How can we build infrastructure that supports high-throughput model training loops and scales to thousands of environments, while keeping every agent’s world consistent, performant, and observable in real time? How to Apply To apply, please email careers@surgehq.ai with a resume and 2-3 sentences describing your interest in Surge. We love personal projects and writings too!
AI 评测地点待确认★★★★★
RL Environments Architect
Surge AI · 地点未写
RL 环境奖励函数奖励作弊仿真遥测
主要职责
- 架构一个模块化的环境框架:清晰的 API、课程学习脚手架、可配置的奖励与终止规则
- 建立质量线:覆盖率指标、不变性检查、对环境输出和 Agent 经验缓冲区做 trace 审计
- 为 episode 回放接入丰富的遥测;缓解奖励作弊(reward hacking)、模式坍塌和可被利用的漏洞
- 与研究员合作,把现实任务转成稳健的仿真,包括合成数据生成器和评测套件
任职要求
- 仿真与系统深度:构建过 RL 环境或模拟器(如自定义物理、多 Agent、工具 API),关注确定性、性能和可观测性
- 数据质量领导力:对设计奖励函数、场景分类体系和 QA 流水线有很强直觉,让信号保持对齐、不漂移
- 构建者心态:能跨研究和工程协作,交付务实、可测试、并能随模型能力演进的环境
这份 JD 的看点:「奖励作弊」(reward hacking)和「模式坍塌」值得先搜一下概念:当 Agent 找到了钻奖励函数漏洞的办法,评测和训练信号就失真了。这份 JD 的思路是把环境当作要治理的软件系统,有覆盖率、不变性检查、审计,几乎就是给环境做「测试工程」。
✅ 你已经有的
- 测试工程思路:覆盖率、不变量检查、trace 审计是测试领域的老概念
- 分布式系统与可观测:Temporal、gRPC、监控体系
- React Flow 等复杂可视化,可用于环境/轨迹的可视化
⚠️ 差距
- RL 环境与奖励设计的背景(需要补)
- 仿真器构建经验
- 地点与薪资未知
🎯 怎么补 / 怎么用
- 先补概念:读一下 OpenAI Gym / Gymnasium 的环境接口(reset / step / reward),自己写一个最小的「工具调用环境」,再故意写一个可被钻空子的奖励函数,观察怎么被利用
原帖:https://surgehq.ai/careers/rl-environments-architect(岗位可能已下线)
查看英文原文(学英语对照用)
The Role As an RL Environments Architect, you’ll design, instrument, and govern the simulated worlds where agents learn — from compact task microcosms to multi-agent, tool-using ecosystems. You’ll define the primitives, reward structures, interfaces, and telemetry that let us stress-test emerging capabilities while keeping training signals faithful, stable, and scalable. Not only will you build environments, you’ll craft standards for data quality and reproducibility across large-scale agent gyms. This is a role for someone who sweats the details of simulation fidelity, thinks in terms of coverage and failure surfaces, and loves turning messy real-world phenomena into learnable curricula. Your work will form the backbone for safe, rapid progress in agentic systems. What You'll Do Architect a modular environment framework with clear APIs, curriculum scaffolds, and configurable reward/termination schemas Establish quality bars: coverage metrics, invariance checks, and trace audits for environment outputs and agent experience buffers Instrument rich telemetry for episode rollouts; mitigating reward hacking, mode collapse, and exploitable loopholes Partner with researchers to translate real-world tasks into robust simulations, including synthetic data generators and evaluation suites What We’re Looking for Simulation & Systems Depth – Experience building RL environments or simulators (e.g., custom physics, multi-agent, tool APIs) with an eye for determinism, performance, and observability Data Quality Leadership – Strong instincts for designing reward functions, scenario taxonomies, and QA pipelines that keep signals aligned and drift-free Builder’s Mindset – Comfort collaborating across research and engineering to ship pragmatic, testable environments that evolve with model capabilities How to Apply To apply, please email talent@surgehq.ai with a resume and 2-3 sentences describing your interest in Surge. We love personal projects and writings too!
AI 评测地点待确认★★★★★
Full Stack Engineer(评测与标注平台)
Surge AI · 地点未写
React/TypeScriptPython标注平台评测工具产品导向
主要职责
- 设计并构建驱动 Surge 平台的核心系统:让人类塑造前沿 AI 模型的工具与界面
- 跨技术栈工作,让数据采集、标注和评测无缝、可规模化,从产品设计决策到后端架构全程负责
- 与产品、研究、运营紧密合作,做出快速、可靠、令人愉悦的工具,帮助全球顶级 AI 实验室训练更安全、更聪明、更强的模型
任职要求
- 用 React/TypeScript 加 Python 或 Ruby on Rails 构建并维护过全栈应用
- 在乎性能、可用性、可维护性的交集,设计简单、可扩展、有韧性的系统
- 能清楚讲解技术概念,喜欢跨职能合作,倾向于快速交付
- 从想法到上线全程负责,快速迭代,写经得起时间考验的代码
加分项 / 其他
- 同公司的 Staff Fullstack Engineer 另有要求:7 年以上生产软件经验;带过范围大且模糊的项目;React/TypeScript + Python/Rails;产品导向;自驱;能向资深工程师解释技术取舍;偏好创业经验;「顶尖院校 CS 学位或同等 STEM 背景」
这份 JD 的看点:这个岗位是「进入 AI 评测公司的门票」:不需要 ML 背景,做的是评测/标注工具本身。JD 很短,说明公司更看你做过什么、能不能快速交付。它的技术栈(React + TS + Python)就是你的主栈。
✅ 你已经有的
- React 18 + TypeScript ★5、Python ★5,正是要求的技术栈
- React Flow / ECharts 的复杂交互,和「标注与评测界面」高度相关
- 做过 AI 测试平台(多租户 + RBAC + 知识库 + 用例库),本身就是「评测/数据管理平台」
- 带 3 人前端团队,技术选型与任务分工
⚠️ 差距
- Rails(可选,Python 也可以)
- 地点与薪资未知
- 没有公开代码/项目可以给他们看(他们明确说喜欢个人项目和代码样例)
🎯 怎么补 / 怎么用
- 最直接的:把 AI 测试平台里能公开的部分做成小 demo(如「AI 生成用例树的编辑器」),录一个 2 分钟视频,投递邮件里带上
原帖:https://surgehq.ai/careers/full-stack-engineer(岗位可能已下线)
查看英文原文(学英语对照用)
The Role As a Fullstack Engineer, you’ll design and build the core systems that power Surge’s platform — the tools and interfaces that let humans shape frontier AI models. You’ll work across the stack to make data collection, labeling, and evaluation seamless and scalable, owning everything from product design decisions to backend architecture.You’ll partner closely with product, research, and operations to create fast, reliable, and delightful tools that help the world’s top AI labs train safer, smarter, and more capable models. This role is ideal for someone who loves moving quickly, cares deeply about product quality, and thrives in high-ownership environments. What We're Looking for Experience building and maintaining fullstack applications using React/TypeScript and Python or Ruby on Rails. You care about the intersection of performance, usability, and maintainability, and you design systems that are simple, scalable, and resilient. You translate technical concepts clearly, enjoy working cross-functionally, and have a bias for shipping. You take ownership from idea to production, iterate fast, and write code that stands the test of time. How to Apply To apply, please email careers@surgehq.ai with your resume and 2–3 sentences describing your interest in Surge. We love personal projects, code samples, or writing too
AI 评测需到岗海外★★★★★
Research Engineer, Machine Learning (Reinforcement Learning)
Anthropic · 英国 · 伦敦
💰 £260,000–£630,000强化学习训练环境评测GPU 集群沙箱代码执行
主要职责
- 团队方向:开发让模型有效使用计算机的系统;通过强化学习推进代码生成;开创面向 LLM 的基础 RL 研究;构建可扩展的 RL 基础设施与训练方法;增强模型推理能力
- 做强化学习的基础研究,通过工具使用创建「Agent 化」模型,处理计算机使用和自主软件生成等开放式任务;提升数学等领域的推理能力;开发内部使用、提效和评测的原型
- 架构并优化核心 RL 基础设施,从清晰的训练抽象,到跨 GPU 集群的分布式实验管理;帮助扩展系统以应对日益复杂的研究流程
- 设计、实现并测试面向 RL Agent 的新训练环境、评测和方法,推动下一代模型的最前沿
- 通过性能分析、优化和基准测试提升整个技术栈的性能
- 实现高效缓存方案并调试分布式系统,加速训练和评测流程
- 与研究和工程团队协作,开发自动化测试框架、设计干净的 API、构建可扩展的基础设施
任职要求
- 精通 Python 及异步/并发编程(如 Trio 框架)
- 有机器学习框架经验(PyTorch、TensorFlow、JAX)
- 有机器学习研究的行业经验
- 能平衡研究探索和工程实现
- 喜欢结对编程;重视代码质量、测试和性能
- 系统设计和沟通能力强;热衷于 AI 的潜在影响,并致力于开发安全、有益的系统
加分项 / 其他
- (强候选人可能具备)熟悉 LLM 架构和训练方法
- 有强化学习技术与环境的经验
- 有虚拟化和沙箱化代码执行环境的经验
- 有 Kubernetes 经验;分布式系统或高性能计算经验;Rust 和/或 C++
- (不需要具备)正式的证书或学历;学术研究经历或论文发表
这份 JD 的看点:这份是看「顶级实验室的 RL 团队到底在做什么」的参考:从训练环境、评测,到 GPU 集群上的分布式实验管理。注意最后一条:不需要学历证书、不需要论文——他们看能力。加分项里有「沙箱化代码执行」和 Kubernetes,这是应用工程师最容易靠近的两块。
✅ 你已经有的
- Python 异步/并发、Kubernetes、分布式系统都是你的能力
- 沙箱与容器化执行:环境管理平台的相关经验
⚠️ 差距
- 机器学习研究背景(PyTorch/JAX、LLM 训练)
- 强化学习经验
- GPU 集群、分布式训练
- 伦敦办公室 25% 到岗;岗位竞争极高
🎯 怎么补 / 怎么用
- 作为「看天花板」的参考即可。想靠近这个方向,最现实的路径是先做 Surge 那类「RL 环境 / 训练基础设施」的工程岗,积累领域经验
查看英文原文(学英语对照用)
About Anthropic Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems. About The Teams Our Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. We've contributed to all Claude models, with significant impacts on the autonomy and coding capabilities of Claude Sonnet 4.5 and Opus 4.5. Our work spans several key areas: Developing systems that enable models to use computers effectively Advancing code generation through reinforcement learning Pioneering fundamental RL research for large language models Building scalable RL infrastructure and training methodologies Enhancing model reasoning capabilities We collaborate closely with Anthropic's alignment and frontier red teams to ensure our systems are both capable and safe. We partner with the applied production training team to bring research innovations into deployed models, and are dedicated to implement our research at scale. Our Reinforcement Learning teams sit at the intersection of cutting-edge research and engineering excellence, with a deep commitment to building high-quality, scalable systems that push the boundaries of what AI can accomplish. About The Role As a Research Engineer within Reinforcement Learning, you will collaborate with a diverse group of researchers and engineers to advance the capabilities and safety of large language models. This role blends research and engineering responsibilities, requiring you to both implement novel approaches and contribute to the research direction. You'll work on fundamental research in reinforcement learning, creating 'agentic' models via tool use for open-ended tasks such as computer use and autonomous software generation, improving reasoning abilities in areas such as mathematics, and developing prototypes for internal use, productivity, and evaluation. Representative Projects Architect and optimize core reinforcement learning infrastructure, from clean training abstractions to distributed experiment management across GPU clusters. Help scale our systems to handle increasingly complex research workflows. Design, implement, and test novel training environments, evaluations, and methodologies for reinforcement learning agents which push the state of the art for the next generation of models. Drive performance improvements across our stack through profiling, optimization, and benchmarking. Implement efficient caching solutions and debug distributed systems to accelerate both training and evaluation workflows. Collaborate across research and engineering teams to develop automated testing frameworks, design clean APIs, and build scalable infrastructure that accelerates AI research. You May Be a Good Fit If You Are proficient in Python and async/concurrent programming with frameworks like Trio Have experience with machine learning frameworks (PyTorch, TensorFlow, JAX) Have industry experience in machine learning research Can balance research exploration with engineering implementation Enjoy pair programming (we love to pair!) Care about code quality, testing, and performance Have strong systems design and communication skills Are passionate about the potential impact of AI and are committed to developing safe and beneficial systems Strong Candidates May Have Familiarity with LLM architectures and training methodologies Experience with reinforcement learning techniques and environments Experience with virtualization and sandboxed code execution environments Experience with Kubernetes Experience with distributed systems or high-performance computing Experience with Rust and/or C++ Strong Candidates Need Not Have Formal certifications or education credentials Academic research experience or publication history Deadline to apply: None. Applications will be reviewed on a rolling basis. The annual compensation range for this role is listed below. For sales roles, the range provided is the role’s On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role. Annual Salary £260,000—£630,000 GBP Logistics Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices. Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this. We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team. Your safety matters to us. To protect yourself from potential scams, remember that Anthropic recruiters only contact you from @anthropic.com email addresses. In some cases, we may partner with vetted recruiting agencies who will identify themselves as working on behalf of Anthropic. Be cautious of emails from other domains. Legitimate Anthropic recruiters will never ask for money, fees, or banking information before your first day. If you're ever unsure about a communication, don't click any links—visit anthropic.com/careers directly for confirmed position openings. How We're Different We believe that the highest-impact AI research will be big science. At Anthropic we work as a single cohesive team on just a few large-scale research efforts. And we value impact — advancing our long-term goals of steerable, trustworthy AI — rather than work on smaller and more specific puzzles. We view AI research as an empirical science, which has as much in common with physics and biology as with traditional efforts in computer science. We're an extremely collaborative group, and we host frequent research discussions to ensure that we are pursuing the highest-impact work at any given time. As such, we greatly value communication skills. The easiest way to understand our research directions is to read our recent research. This research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. Come work with us! Anthropic is a public benefit corporation headquartered in San Francisco. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a lovely office space in which to collaborate with colleagues. Guidance on Candidates' AI Usage: Learn about our policy for using AI in our application process.
AI 评测需到岗海外★★★★★
AI Red Team Lead Engineer
U.S. Bank(美国银行) · 美国 · 夏洛特
💰 $126,000–$149,000AI 红队prompt 注入数据投毒模型反演越狱
主要职责
- 领导 AI 红队行动,对以下对象做对抗性测试:基础模型与自定义模型(LLM、视觉、语音、决策系统);模型部署环境(API、插件、Agent、RAG 流水线);训练、评测与推理流水线;数据摄取、标注和治理控制
- 设计并执行面向 AI 的威胁模拟,对齐真实攻击者、滥用场景和新兴攻击技术(如 prompt 注入、数据投毒、模型反演、越狱、供应链风险)
- 开发并维护定制的 AI 红队工具、框架和自动化,提升测试的规模和可重复性
- 研究新兴 AI 攻击技术、模型漏洞和防御缺口
- 与检测、工程、治理团队合作,支持紫队演练和控制措施验证
- 为 AI 安全标准、测试指南和项目策略做贡献;指导红队工程师并提供技术领导
任职要求
- 本科或同等工作经验;8 年以上相关经验
- 透彻理解信息安全系统、政策与流程;沟通、演示、领导、解决问题和分析能力强
- 有保护 AI/ML 系统(含 LLM 等)的经验;了解 AI 威胁模型与攻击技术(prompt 注入、模型提取、训练数据投毒、推理滥用、幻觉利用)
- 熟悉 AI 平台与工具(模型 API、编排框架、评测流水线)
- 丰富的红队经验,包括对手模拟和多阶段攻击链;能开发概念验证(PoC)利用代码和定制攻击工具
- 熟悉绕过端点与 AI 相关的安全控制(EDR/XDR、API 保护、护栏)
- 有云、容器化和 AI 托管环境的经验;精通 Python、PowerShell、Go、C/C++、Shell 中的一种或多种
- 能把研究转化为可运营的工具;书面与口头沟通出色
这份 JD 的看点:这是「AI 评测的安全一支」的样子:不是评价答案好不好,而是像攻击者一样去打系统。攻击面列得很完整:模型、部署环境(API/插件/Agent/RAG)、训练与推理流水线、数据与标注。可以当作 AI 安全测试的检查清单。
✅ 你已经有的
- Python、Go、容器与云环境
- 反爬突破经历(IP 封锁、验证码、JS 加密签名)说明你有一定的「对抗思维」,但这不等同于安全红队经验
⚠️ 差距
- 信息安全和红队实战(8 年安全经验是硬要求)
- AI 攻击技术的系统知识
- 每周 3 天到岗,美国夏洛特
🎯 怎么补 / 怎么用
- 不建议作为求职目标;但作为学习,可以用 Garak 或 PyRIT 这类开源红队工具对自己的项目扫一遍
原帖:https://www.linkedin.com/jobs/view/ai-red-team-lead-engineer-at-u-s-bank-4409470351(岗位可能已下线)
查看英文原文(学英语对照用)
At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition to life, and each person is unique in their potential. A career with U.S. Bank gives you a wide, ever-growing range of opportunities to discover what makes you thrive at every stage of your career. Try new things, learn new skills and discover what you excel at—all from Day One. Job Description The AI Red Team Lead Engineer leads the execution and evolution of offensive security activities focused on AI/ML systems, platforms, and integrations, in addition to traditional enterprise attack surfaces. Owns the design, execution, and reporting of AI-focused red team operations, adversarial testing, and threat emulation exercises targeting models, data pipelines, AI-enabled applications, and supporting infrastructure. Acts as a senior technical authority and program lead for AI red teaming, partnering with teams to identify, validate, and communicate AI-related risks. Drives maturity through repeatable testing methodologies, automation, custom tooling, and clear articulation of business impact. Key Responsibilities: Lead AI Red Team operations, including adversarial testing of: Foundation and custom models (LLMs, vision, speech, decision systems) Model deployment environments (APIs, plugins, agents, RAG pipelines) Training, evaluation, and inference pipelines Data ingestion, labeling, and governance controls Design and execute AI-specific threat emulation aligned to real-world adversaries, misuse scenarios, and emerging attack techniques (e.g., prompt injection, data poisoning, model inversion, jailbreaks, supply chain risks). Develop and maintain custom AI red team tooling, frameworks, and automation to scale testing and improve repeatability. Perform security research into emerging AI attack techniques, model vulnerabilities, and defensive gaps. Partner with detection, engineering, and governance teams to support purple-team and control validation activities. Contribute to AI security standards, testing guidance, and program strategy. Mentor and provide technical leadership to red team engineers. Required Skills/Experience: Bachelor's degree, or equivalent work experience Eight plus years of relevant experience Thorough understanding of the applicable information security systems, policies, and procedures Effective communication, presentation skills, leadership, problem-solving and analytical skills Proven collaboration and influencing skills Preferred Skills/Experience: Hands-on experience testing or securing AI/ML systems, including LLMs or other model classes Knowledge of AI threat models and attack techniques including, but not limited to, prompt injection, model extraction, training data poisoning, inference abuse, hallucination exploitation Familiarity with AI platforms and tooling (e.g., model APIs, orchestration frameworks, evaluation pipelines) Significant red team experience, including adversary emulation and multi-stage attack chains Proven skill developing proof-of-concept exploits and custom offensive tooling Strong understanding of red team and offensive AI techniques and tooling Expertise defeating or bypassing endpoint and AI-adjacent security controls (EDR/XDR, API protections, guardrails) Experience with cloud, containerized, and AI-hosting environments Proficiency in one or more languages (e.g., Python, PowerShell, Go, C/C++, Shell) Ability to translate research into operational tooling Exceptional written and verbal communication skills This role requires working from a U.S. Bank location three (3) or more days per week. If there’s anything we can do to accommodate a disability during any portion of the application or hiring process, please refer to our disability accommodations for applicants. Benefits: Our approach to benefits and total rewards considers our team members’ whole selves and what may be needed to thrive in and outside work. That's why our benefits are designed to help you and your family boost your health, protect your financial security and give you peace of mind. Our benefits include the following: Healthcare (medical, dental, vision) Basic term and optional term life insurance Short-term and long-term disability Pregnancy disability and parental leave 401(k) and employer-funded retirement plan Paid vacation (from two to five weeks depending on salary grade and tenure) Up to 11 paid holiday opportunities Adoption assistance Sick and Safe Leave accruals of one hour for every 30 worked, up to 80 hours per calendar year unless otherwise provided by law Review our full benefits available by employment status here. U.S. Bank is an equal opportunity employer. We consider all qualified applicants without regard to race, religion, color, sex, national origin, age, sexual orientation, gender identity, disability or veteran status, and other factors protected under applicable law. E-Verify U.S. Bank participates in the U.S. Department of Homeland Security E-Verify program in all facilities located in the United States and certain U.S. territories. The E-Verify program is an Internet-based employment eligibility verification system operated by the U.S. Citizenship and Immigration Services. Learn more about the E-Verify program. The salary range reflects figures based on the primary location, which is listed first. The actual range for the role may differ based on the location of the role. In addition to salary, U.S. Bank offers a comprehensive benefits package, including incentive and recognition programs, equity stock purchase 401(k) contribution and pension (all benefits are subject to eligibility requirements). Pay Range: $126,820.00 - $149,200.00 U.S. Bank will consider qualified applicants with arrest or conviction records for employment. U.S. Bank conducts background checks consistent with applicable local laws, including the Los Angeles County Fair Chance Ordinance and the California Fair Chance Act as well as the San Francisco Fair Chance Ordinance. U.S. Bank is subject to, and conducts background checks consistent with the requirements of Section 19 of the Federal Deposit Insurance Act (FDIA). In addition, certain positions may also be subject to the requirements of FINRA, NMLS registration, Reg Z, Reg G, OFAC, the NFA, the FCPA, the Bank Secrecy Act, the SAFE Act, and/or federal guidelines applicable to an agreement, such as those related to ethics, safety, or operational procedures. Applicants must be able to comply with U.S. Bank policies and procedures including the Code of Ethics and Business Conduct and related workplace conduct and safety policies. Posting may be closed earlier due to high volume of applicants.
AI 评测限定国家★★★★★
Machine Learning Engineer, Global Public Sector
Scale AI · 英国 · 伦敦
💰 $80,000–$120,000Agent harness评测基准红队上下文工程主权 AI
主要职责
- 应用 Agent 研究:主导设计可靠的多步 Agent 系统和长时程推理框架,为国家安全与公共政策场景解决复杂问题
- 系统性评测与红队:开发严谨的基准和评测协议,确保 AI 系统在高风险、非商业环境下安全、无偏见、表现良好
- 模型优化与选型:深入研究模型表现(开放权重和闭源),为细分领域选出最合适的工具,通过上下文工程、RAG 及其他推理时技术进行优化
- 架构 Agent 系统:设计并构建 Agent 架构、harness、工具使用协议和逻辑流,使 LLM 能在复杂工作流中作为可靠的自主 Agent 运行
- 保障可靠性与安全:研究并实现稳健的评测框架,包括面向主权 AI 需求的红队,以及在受监管数据环境下缓解幻觉的策略
- 综合深度研究:构建能自主综合信息、进行长时程推理的 Agent,帮助用户分析海量数据、提取可行动的洞察
- 为细分领域评估并适配模型;在 GPU 受限环境下优化
- 建设评测前沿:创建新的自动化基准,定义公共部门 AI 的成功标准,确保系统满足准确性与主权的最高标准
- 作为技术权威咨询:向公共部门领导者说明新兴 AI 技术的实际边界、安全要求和性能取舍
任职要求
- 工程严谨性:Python 精通,有构建 Agent harness 或 AI 基础设施的经验,写模块化、可扩展、可靠的生产级代码
- 应用研究思维:把理论 AI 概念做成可用原型或产品的经历;能读论文并判断其方法能否用于生产系统
- 评测专长:有 LLM 基准测试、红队或超越标准学术数据集的评测经验
- 高学历(偏好):计算机、数学或相关专业(侧重 ML)硕士或博士,但更看重实际成果和工程卓越
加分项 / 其他
- Agent 系统专家:构建多 Agent 系统的深厚经验,含思维链优化和工具调用可靠性
- 主权 AI 经验:受监管数据环境、本地部署或敏感政府用例
- 推理优化:在有限 GPU 或特定延迟要求下优化模型表现
- 从零到一的心态:适应模糊,享受从头定义研究方向
这份 JD 的看点:它把「Agent 设计」「评测/红队」「模型选型与优化」三件事放进一个岗位,是 Agent 开发和评测两个方向交叉的样本。「读论文并判断能不能落地」这类要求,是应用研究工程师的核心能力。注意这条是一年前的旧帖,参考价值大于投递价值。
✅ 你已经有的
- Python 生产级;Agent harness(nanobot)和 MCP 工具使用协议有实践
- RAG、上下文工程有接触
⚠️ 差距
- 评测/红队的系统性经验
- GPU 受限环境的推理优化
- 公共部门/主权 AI 的行业背景
- 岗位发布时间较早,可能已下线
🎯 怎么补 / 怎么用
- 把「工具调用可靠性」量化:给你的 Agent 列 30 个工具调用场景,统计成功率、参数错误率、无效调用率——这个数据本身就是很好的作品素材
原帖:https://wellfound.com/jobs/2996528-machine-learning-engineer-global-public-sector(岗位可能已下线)
查看英文原文(学英语对照用)
Scale’s mission is to develop reliable AI systems for the world's most important decisions. Our core work consists of: Creating custom AI applications that will impact millions of citizens Generating high-quality training data for national LLMs Upskilling and advisory services to spread the impact of AI Scale is hiring ML Research Engineers to bridge the gap between emerging AI capabilities and mission-critical, real-world impact. In our Global Public Sector (GPS) division, we don’t just implement tools; we conduct applied research to solve the unique challenges of sovereign AI. Your role is to move beyond off-the-shelf implementations. You will lead the research into Agent Design, Reliability, and AI Safety, developing novel system architectures that power high-stakes government applications. You will be the bridge between a research paper and a production-ready system that functions at scale. The Mission Applied Agent Research: Leading the design of reliable, multi-step agentic systems and long-horizon reasoning frameworks that can solve complex problems for national security and public policy. Systemic Evaluation & Red-Teaming: Developing rigorous benchmarks and evaluation protocols to ensure AI systems are safe, unbiased, and performant in high-stakes, non-commercial environments. Model Optimisation & Selection: Conducting deep-dive research into model performance (both open-weight and closed) to identify the best tools for niche domains, optimising them through context engineering, RAG, and other inference-time techniques. What You Will Do Architect Agentic Systems: Design and build agent architectures, the harnesses, tool-use protocols, and logic flows that allow LLMs to function as reliable, autonomous agents in complex workflows. Drive Reliability & Safety: Research and implement robust evaluation frameworks. This includes red-teaming for sovereign AI requirements and developing strategies to mitigate hallucinations in regulated data environments. Synthesise Deep Research: Build agents capable of autonomous information synthesis and long-horizon reasoning, enabling users to analyse massive datasets and extract actionable insights. Optimize for Niche Domains: Evaluate and adapt models for specialised use cases, such as LLM reasoning for low-resource languages, complex OCR tasks, or working in GPU-constrained environments Build Evaluation Frontiers: Create new, automated benchmarks that define what success looks like for AI in the public sector, ensuring our systems meet the highest standards of accuracy and sovereignty. Consult as a Technical Authority: Act as a subject matter expert for public sector leaders, advising on the practical limits, safety requirements, and performance trade-offs of emerging AI technologies. Ideally, You Have Engineering Rigour: Exceptional proficiency in Python and experience building agentic harnesses or AI infrastructure. You write production-ready code that is modular, scalable, and reliable. Applied Research Mindset: A track record of taking theoretical AI concepts and turning them into functional prototypes or products. You know how to read a paper and determine if its methods are actually viable for a production system. Evaluation Expertise: Experience in LLM benchmarking, red-teaming, or building evaluations that go beyond standard academic datasets. Advanced Degree: A Master’s or PhD in Computer Science, Mathematics, or a related field (with a focus on ML) is preferred, but we value demonstrated impact and engineering excellence. Nice to Haves Agentic Systems Expert: Deep experience in building multi-agent systems, including chain-of-thought optimisation and tool-calling reliability. Sovereign AI Experience: Experience working with highly regulated data environments, on-premise deployments, or sensitive government use cases. Inference Optimisation: Knowledge of how to optimise model performance for environments with limited GPU capacity or specific latency requirements. Zero-to-One Mindset: You are comfortable navigating ambiguity and enjoy defining research directions from scratch to solve a specific product or mission need. For those applying based in Qatar: Residency and employment in Qatar requires certain permissions (visa and permits) issued by the Qatari authorities. As part of the application process, candidates will be asked to provide personal information, including residency status and nationality, which is required for visa processing. This information is collected solely for immigration compliance purposes and will be not used as a criterion in any hiring decision unless such use is lawfully permitted. *Visa issuance is at the discretion of the Qatari authorities. If you are successful in your application, you will be required to provide the documentation requested by Scale and the authorities to obtain the necessary permissions for you to live and work in Qatar, and Scale will work with successful candidates to support the visa application process.*
这 27 篇 JD 里反复出现的能力
按英文原文统计,正则已收紧并抽样验证过上下文(比如 Go 只统计作为编程语言出现的,护栏/红队、多 Agent 命中的上下文都看过)。样本只有 27 篇,看趋势,别当结论。
你的位置 & 学习路径
✅ 你的强项(在多份 JD 里都对得上)
- Python 后端 + FastAPI + 微服务/异步(Toptal、ActAI、Mapgenesys)
- Go + K8s/Docker + CI/CD 流水线与发布门禁(Bybit、Surge Backend)
- Temporal 持久化工作流:Agent 长任务、重试、人在回路的基础设施(pst.ag、Surge Backend)
- React/TS 复杂交互 + 可视化:标注/评测/可观测看板(Surge Full Stack、Arize)
- 质量部 + 测试平台背景:评测工程本质是「测非确定性系统」(Dealstitch、EY、Dialpad、Aspire)
- 已有 LLM 接入、RAG、MCP、自研 Agent 框架的实践
⚠️ 缺口(按重要性)
- 评测方法与实操:LLM-as-judge 校准、回归集、步骤级评测
- 可观测/追踪:Langfuse、OpenTelemetry、LangSmith
- Agent 框架实战:LangGraph 的状态、checkpoint、条件路由
- 护栏与安全测试:prompt 注入、OWASP LLM Top 10
- 公开作品:Arize、Surge 都明说看开源/个人项目,你目前没有公开仓库
- 次要:微调、vLLM、Agent 仿真/基准平台
四步学习路径(每步 1–2 周,做完就有作品)
- 评测入门(2 周)用 Promptfoo 或 DeepEval 给你 AI 测试平台里的 RAG 模块写 20–30 条回归用例,接进 CI,指标下降就红灯。对应:Dealstitch、Aspire、Dialpad。
- 可观测 + LangGraph(2 周)用 LangGraph 搭一个 3 步 Agent,接 Langfuse 追踪,对「选工具 / 检索 / 最终答案」分别打分(步骤级评测)。对应:Mapgenesys、Cadence、EY。
- 安全与红队(1 周)读 OWASP Top 10 for LLM Applications,用 Garak 或 PyRIT 对自己的 Agent 扫一遍,写一页报告。对应:EY、U.S. Bank、Dialpad。
- 作品公开(持续)把 nanobot 开源,加上评测与追踪;再做一个「Temporal 持久化 Agent」demo。README 与博客用英文写,顺便练英语。对应:Arize、Surge、Toptal。
投递优先级(我的建议)
- Surge · Backend Engineer:技术最贴(K8s/容器/Temporal/Go),JD 无口语要求。先邮件问清地点与薪资币种。
- Surge · Full Stack Engineer:主栈完全一致,进入评测公司的入口。
- Arize · Open Source AI Engineer:全球远程、薪资明确,前提是先有公开开源作品。
- Toptal · Python Backend + RAG + Agentic:亚洲限定,8 年要求你满足,注意雇佣形式与美国时区重叠。
- Dealstitch · AI Quality Engineer:最清晰的评测岗,但限美国;主要价值是当学习大纲。
术语表
这些词在 JD 里反复出现,先弄懂再读会顺很多。
- Agent / Agentic智能体
- 能自己拆解目标、决定下一步、调用工具并根据结果继续行动的 LLM 系统,区别于「一问一答」的聊天机器人。
- Tool use / Function calling工具调用
- 模型输出结构化调用(函数名 + 参数),由程序去执行并把结果回填给模型。Agent 能「做事」的基础。
- MCP(Model Context Protocol)模型上下文协议
- Anthropic 提出的开放协议,让 Agent 用统一方式发现和调用外部工具/数据源,类似「AI 的 USB 接口」。
- ReAct推理 + 行动
- 让模型交替「思考一步 → 调一次工具 → 看结果 → 再思考」的经典 Agent 模式。
- Reflexion自我反思
- Agent 失败后用文字总结教训,带着这份反思重试,不更新模型权重。
- Orchestrator-workers总控—工人模式
- 一个总控 Agent 拆任务、分派给多个工人 Agent,再汇总结果。
- Multi-agent多 Agent 系统
- 多个各有角色的 Agent 协作、分工或辩论来完成任务。
- Orchestration编排
- 决定 Agent/步骤的执行顺序、分支、重试和状态保存的那一层,LangGraph、Temporal 都在做这件事。
- Durable execution持久化执行
- 任务跑很久、中途会失败,系统能保存状态并从断点恢复。Temporal 的核心能力。
- RAG(Retrieval-Augmented Generation)检索增强生成
- 回答前先从知识库检索相关片段,再连同问题一起交给模型,减少胡编。
- Chunking / Rerank / Hybrid search分块 / 重排 / 混合检索
- 把文档切段入库;先粗召回再用更准的模型重新排序;同时用关键词检索和向量检索。「朴素 RAG」常败在这几步。
- Embedding / Vector DB向量 / 向量库
- 把文本变成数字向量以比较语义相似度;向量库(Pinecone、Qdrant、pgvector…)负责存储和近邻检索。
- Context engineering上下文工程
- 决定每一步把什么信息放进模型的上下文窗口:检索、记忆、工具结果、历史压缩等,比单纯写 prompt 范围更大。
- GraphRAG / Context Graph图检索增强
- 把知识组织成实体与关系的图,再做检索,适合需要多跳推理的问题。
- Guardrails护栏
- 限制模型输入输出的安全层:拦截违规内容、校验格式、限制可调用的工具与权限。
- Human-in-the-loop人在回路
- 关键动作先让人确认或审批,再由系统执行。
- Prompt injection提示注入
- 攻击者把恶意指令藏在用户输入或被检索的文档里,诱使模型偏离原本的指令。间接注入(藏在网页/文档里)更难防。
- Jailbreak越狱
- 用特殊话术绕过模型的安全限制。
- Red teaming红队测试
- 像攻击者一样主动去攻击自己的模型/系统,找出漏洞和失败模式。
- Hallucination幻觉
- 模型自信地编造不存在或不正确的内容。
- Groundedness / Faithfulness依据性 / 忠实度
- 回答是否忠于给定的检索材料,没有编造材料之外的东西。RAG 评测的核心指标。
- Eval(evaluation)评测
- 用一组任务和打分标准,系统性地衡量模型或 Agent 好不好。
- Golden set / Golden dataset黄金数据集
- 人工确认过「标准答案或标准判断」的测试集,用来做基准与回归。
- Rubric评分细则
- 把「好」拆成可逐条打分的标准,供人或 LLM judge 使用。
- LLM-as-a-judge大模型当评委
- 用一个 LLM 给另一个模型的输出打分。必须和人工标注做「校准」(对比一致率),否则分数不可信。
- Regression eval回归评测
- 改了 prompt、模型或检索之后,重新跑同一批用例,确认没有变差。
- Benchmark基准测试
- 公开或内部统一的标准题库,用来横向比较模型。
- Step-level eval步骤级评测
- 不只看最终答案,还给每一步(选哪个工具、检索到什么、推理链)单独打分。Agent 评测的关键区别。
- Reward hacking奖励作弊
- 模型找到钻奖励函数漏洞的办法,拿高分却没有真正完成任务。
- RL environment强化学习环境
- 供 Agent 反复尝试并获得奖励反馈的模拟世界,包含任务、工具、奖励规则。
- SFT / LoRA / DPO / RLHF微调与对齐
- SFT 监督微调;LoRA 只训练少量附加参数的高效微调;DPO、RLHF 用人类偏好数据让模型更符合期望。
- vLLM / Inference serving推理服务
- 把开源模型部署成高吞吐 API 的引擎;关心首 token 时间、吞吐、并发、显存占用。
- Observability / Tracing可观测 / 链路追踪
- 记录 Agent 每一步的输入输出、耗时、token 与成本,出问题时能回放定位。Langfuse、LangSmith、OpenTelemetry。
- ADLC(Agent Development Lifecycle)Agent 开发生命周期
- 规格 → 评测 → 构建 → 测试 → A/B → 灰度 → 监控 → 迭代,把 Agent 当软件来工程化。
- Forward Deployed Engineer前线部署工程师
- 驻场或贴近客户,把产品落地到客户环境的工程师,兼顾部署、开发与客户沟通。