Design and operate production Agentic RAG and LLM agent systems, including adaptive retrieval, multi-hop workflows, agent harness runtimes, memory, tool calling, benchmarking, evaluation, and real-world feedback loops. Develop research-driven prototypes and improve agent intelligence across retrieval efficiency, latency, groundedness, and task success.
Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million people in 100+ countries for our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-asset products. Binance offerings range from trading and finance to education, research, payments, institutional services, Web3 features, and more. We leverage the power of digital assets and blockchain to build an inclusive financial ecosystem to advance the freedom of money and improve financial access for people around the world.
Responsibilities
- Agentic RAG & Engineering: Design and operate next-generation retrieval pipelines — moving beyond static retrieve-once patterns to adaptive, self-correcting, and multi-hop retrieval workflows; architect Agentic RAG systems with dynamic retrieval control, query decomposition, iterative retrieve-reflect-refine loops, and multi-agent retrieval collaboration
- Frontier Harness: Collaborate deeply with researchers and engineers to define and implement model-capability-driven innovations — including context management, long-term memory, subagent and multi-agent architectures, self-evolving agents, and real-word task execution
- Benchmarking & Evaluation: Propose harness-domain and RAG-domain benchmarks and evaluation methodologies; construct benchmark datasets, define annotation strategies, and systematically measure and improve agent intelligence across domains — including retrieval efficiency, latency, groundedness, and task success rate
- Real-world Feedback Loops: Leverage multi-channel user feedback and real-world task data as primary research signals; design experiments and datasets to continuously improve agent and retrieval performance in production scenarios
- Frontier Harness: Collaborate deeply with researchers and engineers to define and implement model-capability-driven innovations — including context management, long-term memory, subagent and multi-agent architectures, self-evolving agents, and real-word task execution
- Benchmarking & Evaluation: Propose harness-domain and RAG-domain benchmarks and evaluation methodologies; construct benchmark datasets, define annotation strategies, and systematically measure and improve agent intelligence across domains — including retrieval efficiency, latency, groundedness, and task success rate
- Real-world Feedback Loops: Leverage multi-channel user feedback and real-world task data as primary research signals; design experiments and datasets to continuously improve agent and retrieval performance in production scenarios
Requirements
- 2-8+ Year hands-on experience with LLM, RAG and AI agent systems in production
- RAG & Agentic RAG Engineering: Hands-on experience building production retrieval pipelines end-to-end — embedding models (BGE, OpenAI, etc.), vector stores (Qdrant, Milvus, Pinecone, Weaviate), hybrid search (keyword + vector), reranking models; deep understanding of chunking strategy, text cleaning, and multimodal data parsing; experience implementing Agentic - RAG patterns — Self-RAG, Corrective RAG, adaptive retrieval, multi-hop decomposition, retrieve-reflect-refine loops
- Agent Harness Engineering — hands-on experience with Agent Harness runtimes (Pi Agent, AgentScope 2.0 or equivalent orchestration frameworks): session recovery, sandbox isolation, middleware/hook systems, multi-tenant runtime, plan/execute loops, and retrieval-grounded tool calling
- LLM & Agent Fundamentals: Deep familiarity with LLM and agent mechanisms — LLM APIs, KV Cache, Agent Loop, Tool Use, Reasoning, Planning, Skills, MCP, Memory, Subagent, Multi-Agent; strong grasp of Prompt Engineering, Context Engineering
- Independent Research Capability: Can analyze ambiguous problems from first principles, generate original ideas, and drive research from 0 to 1; able to rapidly translate ideas into runnable prototypes with tight experiment iteration loops
- Heavy Agent User: Power user of agent products (coding agents, general-purpose agents); agent tools are already integrated into your daily work and life; you have taste and judgment about model behavior
- AI-native Engineering: Proficient in vibe coding — ships fast using AI-assisted workflows across unfamiliar languages, frameworks, and domains; strong learning velocity in software development
- RAG & Agentic RAG Engineering: Hands-on experience building production retrieval pipelines end-to-end — embedding models (BGE, OpenAI, etc.), vector stores (Qdrant, Milvus, Pinecone, Weaviate), hybrid search (keyword + vector), reranking models; deep understanding of chunking strategy, text cleaning, and multimodal data parsing; experience implementing Agentic - RAG patterns — Self-RAG, Corrective RAG, adaptive retrieval, multi-hop decomposition, retrieve-reflect-refine loops
- Agent Harness Engineering — hands-on experience with Agent Harness runtimes (Pi Agent, AgentScope 2.0 or equivalent orchestration frameworks): session recovery, sandbox isolation, middleware/hook systems, multi-tenant runtime, plan/execute loops, and retrieval-grounded tool calling
- LLM & Agent Fundamentals: Deep familiarity with LLM and agent mechanisms — LLM APIs, KV Cache, Agent Loop, Tool Use, Reasoning, Planning, Skills, MCP, Memory, Subagent, Multi-Agent; strong grasp of Prompt Engineering, Context Engineering
- Independent Research Capability: Can analyze ambiguous problems from first principles, generate original ideas, and drive research from 0 to 1; able to rapidly translate ideas into runnable prototypes with tight experiment iteration loops
- Heavy Agent User: Power user of agent products (coding agents, general-purpose agents); agent tools are already integrated into your daily work and life; you have taste and judgment about model behavior
- AI-native Engineering: Proficient in vibe coding — ships fast using AI-assisted workflows across unfamiliar languages, frameworks, and domains; strong learning velocity in software development
Why Binance
• Shape the future with the world’s leading blockchain ecosystem
• Collaborate with world-class talent in a user-centric global organization with a flat structure
• Tackle unique, fast-paced projects with autonomy in an innovative environment
• Thrive in a results-driven workplace with opportunities for career growth and continuous learning
• Competitive salary and company benefits
• Work-from-home arrangement (the arrangement may vary depending on the work nature of the business team)
Binance is committed to being an equal opportunity employer. We believe that having a diverse workforce is fundamental to our success.
By submitting a job application, you confirm that you have read and agree to our Candidate Privacy Notice.
Similar Jobs
Blockchain • Fintech • Software • Web3
Design and execute comprehensive test plans for dApps and smart contracts, perform unit/integration/end-to-end testing, run security audits and vulnerability assessments, automate tests within CI/CD pipelines, monitor quality metrics, debug issues, and collaborate with developers and product managers to ensure reliable DeFi/EVM-compatible applications.
Top Skills:
Automated Testing ToolsCi/CdDappsEthereumEvm-Compatible ChainsHardhatSecurity Analysis ToolsSmart ContractsTest LibrariesTruffle
Blockchain • Fintech • Software • Web3
Lead product strategy and lifecycle for Liquid Staking/Restaking products. Drive development, launch, and growth by collaborating with engineering, design, and marketing. Gather stakeholder and user feedback, perform competitive and market research, and prioritize data-driven product improvements to increase adoption and differentiation.
Top Skills:
AgileBlockchainDefiJIRALiquid RestakingLiquid StakingWeb3
Blockchain • Fintech • Software • Web3
Design, implement, and maintain blockchain infrastructure and dApps. Develop and review smart contracts, optimize consensus and cryptography for scalability and security, troubleshoot performance and reliability issues, and collaborate with product and design teams to deliver features while staying current with blockchain advancements.
Top Skills:
DefiEthereumEvmLiquid StakingRestakingSmart ContractsSolidity
What you need to know about the Bengaluru Tech Scene
Dubbed the "Silicon Valley of India," Bengaluru has emerged as the nation's leading hub for information technology and a go-to destination for startups. Home to tech giants like ISRO, Infosys, Wipro and HAL, the city attracts and cultivates a rich pool of tech talent, supported by numerous educational and research institutions including the Indian Institute of Science, Bangalore Institute of Technology, and the International Institute of Information Technology.

