Sobre esta vaga de AI Ops Engineer na OneMagnify
OneMagnify is an AI native, platform-enabled B2B digital agency operating at the intersection of data, technology, and creativity. We help complex organizations drive measurable business outcomes by building smarter customer experiences and delivering highly integrated solutions across digital, media, and technology. By combining deep industry expertise with advanced analytics and artificial intelligence, we enable our clients to make better decisions, move faster, and compete more effectively in dynamic markets.
We are seeking a motivated and talented AI Operations Engineer with strong cloud infrastructure and AI operations expertise to build, automate, and secure the platform that powers advanced artificial intelligence solutions.
The Impact You’ll Have:
In this role, you will support quantitative analytics and AI engineering teams by designing, automating, and operating end-to-end cloud and AI infrastructure. Responsibilities include managing CI/CD pipelines, containerized microservices, observability platforms, and governance controls to ensure AI applications and models run safely, reliably, and at scale in production environments.
What you’ll do:
- Cloud Platform Engineering: Architect and operate highly available, multi-service AI infrastructure on the cloud, managing the full lifecycle of compute, storage, networking, and security for AI workloads.
- AI-Ops / LLMOps Pipelines: Establish robust MLOps and LLMOps pipelines covering the end-to-end lifecycle of Generative AI tools — including model deployment, version control, prompt and artifact tracking, automated evaluation, and continuous monitoring for performance, drift, and hallucination mitigation.
- CI/CD Automation: Design and maintain automated build, test, and deployment pipelines for full-stack AI applications, ensuring seamless and secure continuous integration and delivery across front-end, back-end, and AI components.
- Infrastructure as Code: Manage cloud infrastructure using Infrastructure as Code (e.g., Terraform, Cloud Build, Kubernetes manifests) to deliver reproducible, auditable, and scalable environments.
- Containerization & Orchestration: Build and operate containerized microservices (Docker/Kubernetes), managing scaling, rolling deployments, resource optimization, and service resilience for AI workloads.
- Observability & Reliability: Implement comprehensive monitoring, logging, tracing, alerting, and SRE practices to ensure platform reliability, availability, and performance of AI applications in production.
- Security & Governance: Embed security across the platform — managing user identities and controlling access rights, safeguarding sensitive credentials, network policies, and data protection — ensuring all AI workloads meet Ford's strict data privacy, security, and compliance standards.
- Developer Enablement: Work closely with AI engineers, software engineers, and analytical modelers to provide self-service tooling, environments, and automated workflows that remove friction from development to production.
What you’ll need:
- Education: Master's or Bachelor's degree in Computer Science, Software Engineering, Cloud Computing, Data Engineering, or a related technical discipline.
- DevOps/Cloud Engineering Experience: 3–5 years of overall experience in cloud engineering, DevOps, or SRE, with at least 1–2 years of dedicated, hands-on experience deploying and operating AI, ML, and Generative AI applications in production.
- AI/MLOps Engineering: Proven track record of implementing CI/CD for AI workloads, containerization (Docker/Kubernetes), and cloud infrastructure management (Terraform, Cloud Build, GKE).
- Automation & Reliability: Demonstrated experience with infrastructure automation, incident response, and building observable, self-healing production systems.
- Cloud Platform: Extensive hands-on experience with a leading cloud provider (e.g., Google Cloud Platform), including Cloud Run, Cloud Build, GKE, GCS, BigQuery, IAM, VPC networking, and Secret Manager.
- DevOps / SRE Practices: Strong proficiency in CI/CD tooling, GitOps, containerization (Docker), orchestration (Kubernetes), Infrastructure as Code (Terraform), and cloud-native monitoring and logging.
- AI-Ops / LLMOps: Proficiency in tools and platforms for model deployment, prompt/model versioning, evaluation, tracing, and monitoring of LLM and GenAI outputs.
- Scripting & Automation: Hands-on experience with a programming/scripting language such as Python, or Bash for automating infrastructure and operational tasks, and for building tooling that serves collaborators.
- Software Engineering Fundamentals: Solid understanding of full-stack application architecture and modern deployment patterns for AI tools, with the ability to integrate front-end, back-end, and AI services reliably.
- Security & Compliance: Familiarity with cloud security best practices, identity and access management, secrets management, and compliance standards in regulated environments.
- Analytics Workflow Understanding: Awareness of the typical workflows of data scientists and modelers (data wrangling, feature engineering, model validation) so you can build reliable platforms and pipelines that serve them.
Future-Ready Skills (Nice to Have):
- Experience in integrated marketing, digital agency, marketing services, or consulting environments preferred.
- Previous exposure to the Banking, Financial Services, or Credit Analytics industries. Experience with Machine Learning engineering and model serving frameworks, or relevant cloud certifications (e.g., Google Cloud Professional DevOps Engineer / Cloud Architect).
Benefits
We offer a comprehensive benefits package including Medical Insurance, PF, Gratuity, paid holidays, and more.
We are an equal opportunity employer
We believe that Innovative ideas and solutions start with unique perspectives. That’s why we’re committed to providing every employee a workplace that’s free of discrimination and intolerance. We’re proud to be an equal opportunity employer and actively search for like-minded people to join our team.