A high-growth Global Semiconductor company pioneering next-generation silicon technology. The company develops sophisticated hardware solutions and a large-scale computational platform to redefine processing performance.
Backed by top-tier investors and a team of industry experts, they are scaling a high-performance hybrid infrastructure that supports massive design and verification workloads at the cutting edge of the industry.
Located in the Central District (adjacent to the train station), the company operates a hybrid model with two days of remote work per week.
Role Overview-
Taking full ownership of the Reliability, Performance, and Automation of a large-scale engineering platform.
Managing and scaling a Hybrid Cloud infrastructure optimized for intensive computational workloads.
Architecting and maintaining production environments using Terraform and Ansible.
Developing self-service automation tools in Python and Bash to empower R&D teams.
Operating and optimizing HPC / Grid scheduling platforms for complex simulation pipelines.
Requirements-
5 years of professional experience in operating large-scale Linux infrastructure – Mandatory.
Proven experience in managing AWS Production environments at scale – Mandatory.
Expert-level proficiency in Terraform, Ansible, Python, and Bash – Mandatory.
Deep understanding of Networking, Storage, and Linux internals.
Hands-on experience supporting HPC or high-performance computing workloads – Mandatory.
Familiarity with Azure or GCP – A significant advantage.
Experience with Cloud Cost Optimization / FinOps – A significant advantage.
Show more
Show less

