Principal Support Engineer – Customer Reliability & Escalations Engineering
hace 1 día
Viladecans
Experteer Overview As a Principal Support Engineer, you will act as a senior technical authority to own and drive high-impact customer escalations. You’ll lead cross-functional incident response, perform deep root cause analysis, and partner with Engineering and Product leadership to implement permanent fixes. You will guide teams through Sev1/Sev2 incidents, shaping platform reliability and customer experience at scale. This role offers strategic influence, mentoring, and direct collaboration with senior stakeholders to solve complex problems. Compensaciones / Beneficios • Own and lead resolution of high-priority Sev1 and Sev2 incidents impacting strategic customers • Lead end-to-end troubleshooting across applications, integrations, cloud infrastructure, databases, and distributed systems • Coordinate cross-functional incident response and technical war rooms with Engineering, Product, Operations, Security, and partners • Conduct deep root cause analysis using logs, metrics, traces, and diagnostics • Partner with Engineering to implement, validate, and communicate permanent corrective actions • Identify recurring issues and drive reliability, observability, automation, and operational efficiency improvements • Serve as trusted technical advisor during executive-level escalations and critical business events • Mentor engineers, influence technical priorities, and promote incident management best practices Responsabilidades • 15+ years in Technical Support Engineering, SRE, Platform Engineering, Application Support, or Production Support • Proven success owning enterprise-critical Sev1/Sev2 incidents through resolution • Experience supporting large-scale SaaS/cloud/enterprise software in customer-facing environments • Strong expertise troubleshooting REST APIs, integrations, authentication services, distributed systems, cloud platforms, databases, and microservices • Hands-on experience with AWS, Azure and/or GCP, Kubernetes, containers, Linux, and modern observability tools (Datadog, Dynatrace, Splunk, Grafana, New Relic, AppDynamics, Elastic) • Ability to conduct complex root cause investigations across multiple domains and stakeholders • Excellent communication skills with engineers, architects, business leaders, Directors, and VPs • Preferred: Software Engineering, DevOps, SRE, CI/CD, scripting (Python, Bash, PowerShell), experience in large-scale cloud/fintech/enterprise environments Requisitos principales • Remote work • career growth opportunities • diverse, inclusive environment