About this Data/AI Engineer (Databricks, AWS, Python, AI)) role at Vanguard
About Vanguard
Founded in 1975, Vanguard is one of the world's leading investment management companies. The firm offers investments, advice, and retirement services to tens of millions of individual investors around the globe—directly, through workplace plans, and through financial intermediaries.
Vanguard India
Vanguard’s office in India is a significant milestone in our global expansion. We are committed to establishing an enduring technology center in Hyderabad, Telangana and are excited to be adding talent who will focus on Artificial Intelligence (AI), mobile, and cloud-based technologies that drive our business outcomes and deliver a world-class experience for our clients.
Role Summary
The Data/AI Engineer is responsible for designing, developing, enhancing, and supporting scalable cloud-based data pipelines, data products, and data-driven applications. This is a hands-on engineering role requiring strong technical expertise in modern data engineering technologies, problem-solving abilities, and a product mindset to deliver reliable, secure, and high-quality data solutions.
The role requires the ability to work independently on assigned initiatives while collaborating closely with Technical Leads, Product Owners, Architects, Analysts, and other Engineering teams. The individual will contribute to the design, development, deployment, enhancement, and operational support of modern data platforms and data products that enable business outcomes and data-driven decision making.
In addition to delivering new capabilities, the Data/AI Engineer will be responsible for supporting existing data products and applications, performing production support activities, troubleshooting incidents, implementing enhancements, and driving continuous improvement. The individual is expected to take ownership of assigned deliverables from development through production support while seeking guidance from the Technical Lead for complex technical, architectural, or cross-functional decisions.
Responsibilities
- Build and optimize data ingestion, transformation, integration, and storage solutions supporting analytical and operational workloads.
- Design, develop, test, and maintain scalable batch and real-time data pipelines, data products, and cloud-based data solutions.
- Own assigned deliverables from design through deployment, ensuring quality, reliability, performance, security, and maintainability.
- Collaborate with Technical Leads, Product Owners, Architects, Analysts, and Engineering teams to translate business requirements into technical solutions.
- Implement data quality controls, monitoring, observability, automation, and CI/CD practices to improve platform reliability and engineering productivity.
- Support and enhance business-critical data products, applications, and platforms through defect resolution, performance improvements, and continuous modernization initiatives.
- Troubleshoot production issues, perform root-cause analysis, implement preventive actions, and contribute to operational excellence.
- Participate in application-support activities, release support, and production operations, including flexible working hours when required to support critical deliverables, production incidents, business commitments, or global stakeholders.
- Contribute to technical documentation, reusable components, engineering standards, and AI-assisted development practices that improve quality and delivery efficiency.
- Proactively identify risks, dependencies, and improvement opportunities while working independently with guidance from the Technical Lead when required.
Qualifications and Skills
- Bachelor's degree in Computer Science, Information Technology, Engineering, Computer Applications, or a related discipline from a recognized institution.
- Minimum of 5 years of relevant experience with the majority of experience focused on Data Engineering.
- Hands-on experience designing, developing, and supporting scalable data pipelines, ETL/ELT frameworks, and cloud-based data solutions.
- Experience delivering data engineering solutions in cloud environments and Agile product teams.
- Demonstrated ability to independently deliver assigned workstreams while collaborating effectively with Technical Leads and Architects.
- Experience supporting production data platforms, data products, and business-critical applications.
- Strong understanding of software development lifecycle, production support, operational excellence, and continuous improvement practices.
Must have Skills
- Cloud Platforms- Hands-on experience with AWS data services such as S3, Glue, Lambda, Redshift, Athena, Step Functions, SNS/SQS, IAM, CloudWatch, and Kinesis, as relevant to the assigned solutions.
- Data Engineering- Strong experience with scalable ETL/ELT pipelines, data integration, batch processing, data transformation, and modern data-platform patterns.
- Programming- Strong hands-on proficiency in Python, PySpark, Pandas, and Advanced SQL, including writing maintainable code and optimizing data-processing workloads.
- Data Platforms- Working experience with Redshift, Databricks, and Spark, including incremental processing, data quality, and performance optimization.
- Orchestration and Streaming- Practical experience with Airflow and Kafka, including scheduling, retries, dependency handling, monitoring, and failure recovery.
- DevOps and Infrastructure- Working knowledge of Docker, Terraform, GitHub Actions, Code Pipeline, ECS, ECR, and CI/CD practices.
- Analytics and Integration- Experience integrating data through REST APIs and supporting datasets consumed by analytics and visualization solutions such as Power BI
- Engineering Practices- Strong understanding of Git, code reviews, automated testing, observability, data quality, security, governance, and production-support practices.
- Core GenAI- Hands on / Familiarity with Prompt & context engineering, RAG, agentic workflows, tool/function calling, structured outputs (Bedrock, LangChain/LangGraph, MCP)
- AWS AI/ML Platform- Hands on / Familiarity of Bedrock (Knowledge Bases, Guardrails, Agents), SageMaker, Lambda, S3, OpenSearch Serverless vectors, Step Functions, IaC (CDK/CloudFormation)
Location
This role is based in Hyderabad, Telangana at Vanguard India. Only qualified external applicants will be considered.
Our mission
Vanguard adheres to a simple purpose: To take a stand for all investors, to treat them fairly, and to give them the best chance for investment success.
Our commitment to you
Vanguard takes the same long-term view of your success—at work and in life—with Benefits and Rewards packages that reflect what you care about, throughout all the phases and stages of your life. Our Total Rewards programs provide you and your loved ones with wellness support for key areas in your life:
Financial wellness
We're committed to enabling your financial success and provide competitive offers and programs.
Physical wellness
We're committed to providing benefits that support your physical health and wellness.
Personal wellness
We're committed to providing resources that help support the full scope of your life.
How we work
Vanguard has implemented a hybrid working model for most of our employees (crew members), designed to capture the benefits of enhanced flexibility while enabling in-person learning, collaboration, and connection. We believe our mission-driven and highly collaborative culture is a critical enabler to support long-term client outcomes and enrich the employee experience.