About this Software Engineer, Production role at Cxnpl
The Role
You will own production issues on a live banking platform, end to end. When an incident lands with you, you find the root cause in the code, the data, or the infrastructure. Then you fix it, or you work with our senior engineers to ship the fix.
You will work across the full platform: payments, lending, deposits, the core ledger, and the services that connect them. Few roles give you this breadth this early in your career.
You will also use AI tools every day to diagnose issues, trace problems across services, and draft fixes. Over time, you will help us build the automation that takes over routine work. The goal is a support function that runs itself, and you will help design it.
What you will do
- Own Medium and Low impact incidents from assignment through to resolution, within agreed SLAs.
- Investigate issues across services, logs, databases, and code to find the true root cause.
- Write fixes where you can. Work with the Change team when a significant hotfix is required.
- Raise problem records for residual impact and recurring incidents. Drive the investigation through to a permanent fix.
- Coordinate Post Incident Reviews for Critical and High priority incidents, with the Engineering Managers who author them.
- Convert PIR and investigation actions into tracked problem tickets in Helix, and follow them through to closure.
- Use AI coding and analysis tools to diagnose and resolve issues faster.
- Turn manual runbooks into scripts, tools, and automated workflows.
- Help design and build the systems that automate triage, diagnosis, and resolution of routine incidents.
- Work closely with our Operations team to give our clients and their customers consistent coverage, care, and support.
- Find and fix issues in production before they become incidents.
Who you are
- You have 2–4 years of experience as a software engineer, ideally with some time in production support or on-call.
- You are strong in at least one language, such as TypeScript, Java, Kotlin, or Python.
- You read and debug code that you did not write, and you are comfortable in a large codebase.
- You know SQL well enough to investigate data issues directly.
- You understand how distributed systems fail, for example timeouts, retries, race conditions, and partial failures.
- You already use AI tools in your work, and you want to push them further.
- You want to understand how the whole platform works, not just one service.
- You stay calm under pressure and communicate clearly with engineers, operations, and stakeholders during incidents.
- You fix root causes, not symptoms.