Sobre esta vaga de Site Reliability Engineer na Binance
Responsibilities
- Frontier Harness: Collaborate deeply with researchers and engineers to define and implement model-capability-driven innovations — including context management, long-term memory, subagent and multi-agent architectures, self-evolving agents, and real-word task execution
- Benchmarking & Evaluation: Propose harness-domain and RAG-domain benchmarks and evaluation methodologies; construct benchmark datasets, define annotation strategies, and systematically measure and improve agent intelligence across domains — including retrieval efficiency, latency, groundedness, and task success rate
- Real-world Feedback Loops: Leverage multi-channel user feedback and real-world task data as primary research signals; design experiments and datasets to continuously improve agent and retrieval performance in production scenarios
Requirements
- RAG & Agentic RAG Engineering: Hands-on experience building production retrieval pipelines end-to-end — embedding models (BGE, OpenAI, etc.), vector stores (Qdrant, Milvus, Pinecone, Weaviate), hybrid search (keyword + vector), reranking models; deep understanding of chunking strategy, text cleaning, and multimodal data parsing; experience implementing Agentic - RAG patterns — Self-RAG, Corrective RAG, adaptive retrieval, multi-hop decomposition, retrieve-reflect-refine loops
- Agent Harness Engineering — hands-on experience with Agent Harness runtimes (Pi Agent, AgentScope 2.0 or equivalent orchestration frameworks): session recovery, sandbox isolation, middleware/hook systems, multi-tenant runtime, plan/execute loops, and retrieval-grounded tool calling
- LLM & Agent Fundamentals: Deep familiarity with LLM and agent mechanisms — LLM APIs, KV Cache, Agent Loop, Tool Use, Reasoning, Planning, Skills, MCP, Memory, Subagent, Multi-Agent; strong grasp of Prompt Engineering, Context Engineering
- Independent Research Capability: Can analyze ambiguous problems from first principles, generate original ideas, and drive research from 0 to 1; able to rapidly translate ideas into runnable prototypes with tight experiment iteration loops
- Heavy Agent User: Power user of agent products (coding agents, general-purpose agents); agent tools are already integrated into your daily work and life; you have taste and judgment about model behavior
- AI-native Engineering: Proficient in vibe coding — ships fast using AI-assisted workflows across unfamiliar languages, frameworks, and domains; strong learning velocity in software development