About Kaseya
Kaseya is the leading provider of AI-powered IT management and cybersecurity software, serving Managed Service Providers (MSPs) and internal IT organizations worldwide. Our comprehensive platform helps organizations efficiently manage, secure, and automate their IT environments, driving operational efficiency and long-term business success.
Backed by Insight Partners, a leading global software investor, Kaseya has experienced sustained double-digit growth and continues to expand its global footprint. Today, Kaseya supports customers in more than 20 countries and manages over 15 million endpoints worldwide.
Founded in 2000, Kaseya has built a culture centered around innovation, accountability, and results. We are a high-growth, high-performance organization that values individuals who are driven, adaptable, and committed to delivering exceptional outcomes for our customers and teammates alike.
At Kaseya, success comes from embracing challenges, moving with urgency, and continuously raising the bar.
About the Role
- We are looking for an experienced Senior DevOps Engineer with deep expertise in Linux Administration to join our Backup Platform Engineering team.
- In this role, you will own the reliability, scalability, and performance of large-scale Linux infrastructure that powers our next-generation backup and disaster recovery platform.
- You will work closely with software engineering, SRE, and platform teams to automate infrastructure, improve operational excellence, and ensure highly available production environments.
Required Skills:
- 8–12 years of DevOps / SRE experience, with at least 3 years managing large-scale infrastructure
- Kubernetes — cluster operations, resource management, custom controllers, multi-tenant workload isolation
- Infrastructure as Code — Terraform or Pulumi at production scale; versioned, modular, reusable
- CI/CD pipeline ownership — designing and maintaining pipelines (GitHub Actions, Jenkins, ArgoCD or equivalent)
- Observability stack — metrics, logs, and traces in production (Prometheus, Grafana, Datadog or equivalent); defining SLOs/SLAs, not just dashboards
- Incident management at scale — structured on-call, alert triage, runbooks, post-mortems; experience reducing alert noise (1K+ alerts/month environment)
- Networking fundamentals — DNS, load balancing, firewalls, VPC/overlay networks in hybrid environments
- Security & compliance mindset — secrets management (Vault), RBAC, image scanning, audit logging; critical for a backup product handling customer data
- Scripting proficiency — Go or Python for automation; shell scripting for ops tooling
- Linux systems depth — performance tuning, kernel parameters, storage I/O, process management at scale
Desired Skills:
- OpenStack operations — managing Nova, Swift, Neutron, Cinder at scale
- Multi-cloud abstraction — managing workloads across AWS, GCP, Azure and private cloud with consistent tooling
- Large-scale infrastructure (5K+ nodes) — capacity planning, hardware lifecycle, rack-level failure domains
- Cost optimization / FinOps — cloud spend analysis, rightsizing, storage tiering strategies
- Chaos engineering — fault injection, game days, resilience testing (Chaos Monkey, Litmus)
- Bare metal provisioning — PXE boot, IPMI, automated OS provisioning at scale (Ironic, MaaS)
- Backup/DR domain awareness — understanding RPO/RTO, storage replication, data protection pipelines
- Go proficiency — reading and debugging Go services, contributing to internal tooling
Additional information
Kaseya provides equal employment opportunity to all employees and applicants without regard to race, religion, age, ancestry, gender, sex, sexual orientation, national origin, citizenship status, physical or mental disability, veteran status, marital status, or any other characteristic protected by applicable law.