Jobs Companies Resilientco Sr Platform/ Infrastructure Engineer

Über diese Sr Platform/ Infrastructure Engineer Stelle bei Resilientco

Resilientco · Remote · Argentina

We are seeking a senior Sr Platform/Infrastructure Engineer to strengthen our platform team and drive cloud-native infrastructure initiatives. This role focuses on deploying and maintaining Kubernetes services, integrating monitoring and storage platforms, and troubleshooting distributed systems to ensure resilient, scalable operations.

You will work with Python-driven tooling, Prometheus-based monitoring, Ceph-backed storage, and public cloud environments (AWS and Azure) to modernize and operate our platform. This is an opportunity to shape platform reliability and performance in a hands-on engineering role.

Responsibilities

  • Design, deploy, and maintain production Kubernetes clusters and related services.
  • Build and maintain automation and tooling using Python to support platform operations.
  • Integrate and operate Prometheus for monitoring, alerting, and observability.
  • Deploy and manage Ceph storage solutions for distributed workloads.
  • Support platform modernization initiatives and migrate services to cloud-native patterns.
  • Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers.
  • Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
  • Document platform designs, runbooks, and operational procedures.
  • Participate in on-call rotations and incident response to maintain platform availability.

Requirements

  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
  • Proven experience deploying and operating Kubernetes in production.
  • Strong Python skills for automation, tooling, and operational scripts.
  • Experience implementing and operating Prometheus-based monitoring and alerting.
  • Hands-on experience with Ceph or similar distributed storage systems.
  • Cloud experience with AWS and Azure (designing, deploying, and operating services).
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
  • Experience collaborating across teams to deliver platform improvements and migrations.

Nice to Have

  • Experience with OpenSearch.
  • Proficiency with Bash scripting.
  • Familiarity with Java-based services.
  • Experience with Fluent Bit for log collection.
  • Experience working with PostgreSQL.

Engagement & Logistics

  • Engagement Length: 12 months or more.
  • Time Zone: PST - 8:00 AM - 5:00 PM
  • Holiday Calendar: Client Holidays (USA – Mandatory)
  • Laptop: BYOD.
  • Overtime Required: No.



Selection process

  1. Meeting with Resilient Co. team.
  2. Technical interview
  3. Client (2 interviews - Manager + Technical panel)
Bereit, sich bei Resilientco zu bewerben?
Bei Resilientco bewerben

Ähnliche Jobs

Registrieren für Vorschläge, die auf die von Ihnen geöffneten Jobs und gespeicherten Suchen zugeschnitten sind.

Mehr Jobs bei Resilientco

Alle Jobs bei Resilientco ansehen →

Jetzt bewerben
🤖

Moment — langsam

JobsRadar wurde für echte Menschen gebaut, die eine schwere Zeit bei der Jobsuche haben — nicht für automatisierte Anfragen. Sie klicken viel zu schnell und sind jetzt vorübergehend blockiert.

Kommen Sie später wieder. Wenn Sie wirklich auf Jobsuche sind, stehen wir hinter Ihnen — verhalten Sie sich einfach wie ein Mensch.

Catch your next role the second it’s posted.

Create a free account and we’ll watch the boards for you — the instant a job matches your search, it lands in your inbox or Telegram. No digging, no refreshing.

Create free account

Free forever · takes 30 seconds · already have one?

Verschaffe dir einen Vorsprung bei der Jobsuche.

Tritt unserem Telegram-Kanal bei für das, was dir hilft, die Stelle zu bekommen — Gehaltsbenchmarks, den wöchentlichen Marktpuls und neue Feature-Drops. Kein Spam, nur Signal.

Dem Kanal beitreten — kostenlos