## Senior Site Reliability Engineer, Efficiency & Performance
### About the Role
The Efficiency & Performance team focuses on improving the reliability, scalability, performance, and cost efficiency of the ThousandEyes cloud platform across AWS environments. The team works closely with engineering, platform, finance, product, and leadership to drive infrastructure optimization, cloud cost governance, performance improvements, and operational excellence.
In this Senior Site Reliability Engineer role, you will own AWS cost optimization and efficiency initiatives and help teams make data-driven decisions about cost, performance, and scalability. The position combines site reliability engineering with FinOps practices to reduce waste, improve utilization, and support ongoing platform growth.
You will also build the automation, dashboards, reports, and guardrails needed to improve visibility into cloud spend and platform health. Success in the role depends on influencing engineering teams, supporting reviews and stakeholder updates, and balancing reliability, scalability, performance, and cost efficiency in platform decisions.
### Key Responsibilities
– Lead AWS cost optimization and efficiency initiatives across ThousandEyes infrastructure.
– Analyze cloud usage across compute, storage, databases, observability, networking, and data platform workloads to identify savings opportunities.
– Work with engineering teams to improve application and infrastructure performance while reducing cloud waste.
– Identify and reduce underutilized infrastructure, idle resources, over-provisioned workloads, and inefficient service usage.
– Improve resource utilization through autoscaling, right-sizing, and capacity planning.
– Build automation, dashboards, reports, and guardrails that improve cost governance and operational visibility.
– Support cost reviews, OKR tracking, leadership updates, and stakeholder communications.
– Partner with product, finance, engineering, and platform teams to align optimization work with business priorities.
– Improve observability, alerting, and monitoring for cost, performance, and platform health indicators.
– Mentor engineers and lead cross-functional initiatives through to completion.
– Promote cost-aware architecture, reliable design patterns, and performance-efficient implementation practices.
### Required Skills
– Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience.
– 8–12 years of relevant experience in SRE, cloud infrastructure, platform engineering, DevOps, production engineering, or performance engineering.
– Practical experience in cloud cost optimization and FinOps, including AWS billing analysis, cost allocation, tagging, budget tracking, forecasting, and cost governance.
– Experience with performance analysis, capacity planning, infrastructure optimization, and reliability improvements for large-scale cloud platforms.
– Experience with infrastructure as code and automation tools.
– Strong scripting or programming skills in Python, Go, Shell, or similar languages.
– Experience with observability and monitoring platforms.
– Strong Linux systems knowledge and understanding of distributed systems.
– Experience in incident management, production support, reliability engineering, and operational excellence.
– Ability to analyze large-scale infrastructure, performance, and cost data and turn findings into actionable recommendations.
– Strong communication skills for presenting technical, performance, and cost insights to engineering, finance, leadership, and cross-functional stakeholders.
– Experience working in Agile/Scrum environments and managing priorities across multiple teams.
### Preferred Skills
– Experience supporting or optimizing large-scale SaaS platforms in AWS.
– FinOps certification or equivalent cloud financial management experience.
– Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle optimization, workload right-sizing, and commitment planning.
– Experience with AWS Cost Explorer, AWS CUR, Cloud-ability, Cloud-Health, or similar cost management tools.
– Experience with performance tuning of cloud infrastructure, distributed systems, databases, data pipelines, and high-scale services.
– Experience with EMR, Kafka, OpenSearch, RDS, Airflow, Spark, Redis, Cassandra, or similar data platform components.
– Experience driving cross-team cost optimization programs, efficiency OKRs, governance reviews, executive reporting, and measurable savings outcomes.
– Ability to influence engineering teams toward cost-conscious architecture, performance-aware design, and operational best practices.
– Exposure to AI/ML infrastructure cost optimization, capacity planning, GPU/accelerator cost governance, or AI-driven efficiency tooling.
### Cloud Platforms & Technologies
– **Cloud Providers:** AWS
– **AWS Services:** EC2, S3, RDS, EMR, Lambda, CloudWatch, OpenSearch, ElastiCache, IAM, VPC
– **FinOps / Cost Management Tools:** AWS Cost Explorer, AWS CUR, Cloud-ability, Cloud-Health
– **AWS Cost Optimization Mechanisms:** AWS Savings Plans, Reserved Instances, Spot, Graviton
– **Programming Languages:** Python, Go, Shell
– **Infrastructure as Code / Automation:** Terraform, CloudFormation, Puppet, Ansible
– **Monitoring / Observability:** ThousandEyes, Prometheus, Grafana, Splunk, Datadog
– **Data / Platform Technologies:** Kafka, Airflow, Spark, Redis, Cassandra
– **Operating Systems & Systems Concepts:** Linux, distributed systems
– **AI / ML Infrastructure:** GPU, accelerators
### FinOps Responsibilities
– Improve cloud cost visibility, governance, tagging hygiene, allocation, showback/chargeback, budgeting, forecasting, and anomaly detection.
– Analyze cloud usage across compute, storage, databases, observability, networking, and data platform workloads to identify savings.
– Reduce underutilized infrastructure, idle resources, over-provisioned workloads, and inefficient service usage.
– Improve resource utilization through autoscaling, right-sizing, and capacity planning.
– Build automation, dashboards, reports, and guardrails to support cost governance and operational visibility.
– Support cost reviews, OKR tracking, leadership updates, and stakeholder communications.
– Align cost optimization work with business priorities across engineering, platform, finance, and product teams.
### Why You Might Be Interested
This role offers direct ownership of AWS cost optimization and platform efficiency. It combines SRE, FinOps, and cross-functional leadership, with responsibility for improving reliability, scalability, performance, and cost efficiency. You would work closely with engineering, finance, product, and leadership, and your work would influence both technical decisions and business outcomes.

