About this Principal Software Engineer role at Cleo (India)
About the Role
We are looking for a Principal Software Engineer to shape how our AWS-based SaaS products are built, deployed, and operated. This is a hands-on technical leadership role for an engineer with prior software development experience and deep expertise in AWS, infrastructure as code, CI/CD, automation, and observability.
You will partner with application, platform, product, and security teams to improve the full software delivery lifecycle. You will modernize existing applications by reducing unnecessary coupling and enabling components to be built, tested, released, and deployed independently where the architecture supports it. You will also set technical direction across teams, lead complex modernization initiatives, and mentor engineers in the Bangalore organization.
What You Will Do
DevOps Strategy and Technical Leadership
- Define DevOps and platform engineering standards, reference architectures, delivery patterns, and operational practices.
- Lead cross-team initiatives that improve delivery speed, reliability, security, and maintainability.
- Make and communicate technical decisions that balance modernization, customer impact, and service continuity.
- Mentor engineers, lead knowledge-sharing sessions, and contribute to hiring and technical interviews.
Application Modernization and Software Engineering
- Apply software development experience to understand application architecture, dependencies, build systems, runtime behavior, and operational risks.
- Assess existing applications and reduce tight coupling between components, services, shared libraries, and deployment units where practical.
- Enable components to be built, tested, versioned, and deployed independently
- Improve configuration, secrets handling, health checks, logging, metrics, and graceful startup and shutdown behavior.
- Build automation and engineering tools using a programming or scripting language such as Python, Go, Groovy, or shell.
CI/CD and Developer Experience
- Design, implement, and operate CI/CD platforms using GitHub Actions and Jenkins.
- Build and maintain reusable workflows, pipeline templates, and shared components using pipeline-as-code.
- Create reliable workflows for builds, automated tests, code quality checks, container creation, artifact management, deployment, and environment promotion.
- Integrate appropriate security and quality checks, including dependency and container scanning and secrets detection.
- Manage workflow permissions and credentials using least privilege and short-lived credentials or federated identity where supported.
- Improve build and deployment reliability through caching, parallel execution, consistent toolchains, and clear failure diagnostics.
- Modernize existing Jenkins pipelines and integrate them with GitHub repositories and Actions where appropriate.
- Improve developer experience with self-service workflows, standard templates, documentation, and actionable pipeline feedback.
- Comfortable with uncertainty and ambiguous requirements, and can adapt quickly as the landscape shifts.
AWS, Kubernetes, and Infrastructure as Code
- Architect, operate, and improve AWS environments, with strong hands-on experience in services such as EKS, ECS, EC2, RDS, S3, Route 53, IAM, ACM, ALB, Lambda, VPC etc.
- Design and operate multi-account AWS environments, including networking, CIDR allocation, routing, VPN connectivity, DNS, certificates, cross-account access, and security controls.
- Build and maintain reusable infrastructure as code using Terraform; use CloudFormation where appropriate.
- Establish practices for IaC module design, environment configuration, review, testing, change planning, drift detection, and state management.
- Use Docker, Kubernetes, Helm and GitOps practices to deliver consistent application deployments.
AI-Augmented Engineering and Operations
- Lead adoption of AI coding assistants and agents across platform work, such as authoring Terraform modules, pipeline templates, and Helm charts, and speeding up Jenkins-to-GitHub-Actions migrations. Keep guardrails in place: code review, policy-as-code, automated tests, and plan/apply approvals.
- Build AI-assisted steps into CI/CD, such as automated PR review, build and test failure triage, flaky-test detection, and plain-language failure summaries for developers.
- Provide platform support for teams shipping AI features. This covers secure access to model APIs (e.g., Amazon Bedrock), GPU or inference workloads on EKS where needed, prompt and model versioning, evaluation gates in pipelines, and cost visibility for AI usage.
- Define governance for AI agents that act on infrastructure. That means least-privilege, short-lived credentials for agents, audit trails, human approval for production changes, and protection against data leakage and prompt injection.
- Measure the impact of AI adoption with delivery and reliability metrics (e.g., DORA), and share what works through the knowledge-sharing sessions already in the role.
What You Bring
- Significant experience in DevOps, Platform Engineering, or a closely related discipline.
- Prior hands-on software development experience and the ability to work effectively with application code, build systems, and service architecture.
- Strong experience modernizing applications and reducing dependencies that unnecessarily constrain build and deployment practices.
- Deep hands-on expertise with AWS tech stacks.
- Strong knowledge of infrastructure-as-code, particularly Terraform, including reusable modules and environment-specific change management.
- Hands-on experience designing and operating CI/CD platforms, including GitHub Actions. Experience with Jenkins or similar systems is also expected.
- Practical experience with Docker, Helm, GitOps, Kubernetes, automated testing, artifact management, and deployment strategies.
- Strong programming or automation skills in at least one of Python, Go, Groovy, or shell scripting.
- Solid understanding of cloud networking, VPCs, routing, load balancing, DNS, VPNs, IAM, TLS certificates, and multi-account architectures.
- Experience with observability and incident-management tools such as Datadog, Prometheus, Grafana, CloudWatch, Alertmanager, or Opsgenie.
- Strong communication skills and the ability to explain technical decisions and trade-offs clearly to technical and business stakeholders.
- Hands-on experience using AI coding assistants or agents in daily engineering work, with sound judgment about when output needs verification and how to review it.
- Understanding of the security and operational risks of AI tooling in delivery pipelines and cloud environments.
- Preferred: experience deploying or operating AI/LLM-powered services on AWS, or building AI-driven automation for CI/CD or incident response.
Preferred Qualifications
- 12+ years of relevant professional experience.
- B.E., B.Tech., M.Tech., MCA, or equivalent qualification in Computer Science, Information Technology, or a related discipline.
Cleo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, color, family or medical care leave, gender identity or expression, genetic information, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran status, race, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable laws, regulations and ordinances.