Gauss Labs

Senior Site Reliability Engineer (KR) at Gauss Labs (Yeoksam, Seoul, South Korea)

Gauss Labs· Yeoksam, Seoul, South Korea·

Role details

Work type
Hybrid
Employment
Full-Time

Job description

Gauss Labs is an industrial AI company on a mission to revolutionize manufacturing with AI, starting with the semiconductor sector. Panoptes is an AI-based virtual metrology solution deployed in high-volume manufacturing fabs, helping customers improve yield, reduce costs, and accelerate production. Our software runs in our customers' own managed environments, and we're seeking a Site Reliability Engineer to own the reliability of the infrastructure and platform that Panoptes runs on. You will keep the platform available, performant, and scalable; own monitoring, alerting, incident first-response, and the on-call rotation; and build the automation and observability that let engineering teams operate their services safely.

Responsibilities

  • Platform reliability and operations: Own platform-layer reliability across both environments. In our internal cloud environment: full ownership — cluster health, resource management (CPU/memory/OOM), scheduling, autoscaling, Kubernetes/EKS lifecycle. In the customer environment: operate directly at the application-namespace level and for the customer-controlled cluster/node layer, diagnose and clearly communicate what's needed, and operate the platform within their setup, decisions, and constraints.
  • Monitoring and Alerting: Build and maintain robust monitoring and alerting for the infrastructure and platform layer to proactively identify and resolve issues before they impact the platform.
  • Incident Response: Own incident first-response for the platform layer and participate in the on-call rotation to minimize downtime and restore service quickly.
  • Automation: Develop automation tools and scripts to streamline operations, reduce manual effort, and enable engineering teams to operate their own services safely.
  • Capacity Planning: Forecast resource needs, optimize resource utilization, and ensure the platform infrastructure can handle increasing workloads.
  • Deployment infrastructure: Build and maintain CI/CD pipelines and deployment infrastructure for the platform.
  • Continuous Improvement: Drive a culture of continuous improvement by identifying opportunities to enhance platform reliability, performance, and efficiency.

Basic Qualifications

  • Bachelor's degree in computer science, engineering, or a related discipline
  • 5+ years of industry experience as a Site Reliability Engineer or in platform/infrastructure engineering
  • Hands-on experience operating Kubernetes in production (EKS preferred): cluster lifecycle, scheduling, autoscaling, resource management
  • Experience with cloud platforms (AWS preferred) and containerization technologies (Docker, Kubernetes)
  • Experience with observability and alerting tools (Prometheus, Grafana, ElasticSearch, Jaeger)
  • Experience with scripting languages (Python, Bash)
  • Working knowledge of GitHub, GitHub Actions, and CI/CD concepts
  • Strong problem-solving and troubleshooting skills
  • Working proficiency in English for internal documentation and technical coordination

Preferred Qualifications

  • Knowledge of AI/ML infrastructure and workloads.
  • Knowledge of database technologies (MongoDB, PostgreSQL)
  • Experience operating software in customer-managed (on-prem or customer-cloud) environments
  • Exposure to manufacturing, semiconductor, or enterprise B2B customer environments

[Interview process] Application reivew - Phone interview - Virtual onsite interview - VP interview/Core Value interview - CEO interview

Why work at Gauss Labs

  • Culture highlights: Gauss Labs emphasizes “normalizing AI” and values ownership, fearless experimentation, and collective learning. The company is described as having a balanced and inspiring leadership.
  • Remote/hybrid/office policy:
    • Palo Alto: Hybrid working environment
    • Seoul: Hybrid working environment
    • Vancouver: Full remote working environment
    • Generous PTO, flexible hours, and remote work options
  • Notable perks: Comprehensive health insurance (medical, vision, dental) for employees and dependents; 401(k) with company matching in the US; group insurance and private health plan in Korea; robust savings plans; career development programs; team-building events.
  • Engineering culture: High proportion of technical staff (38% technical, 14% research), with strong talent drawn from top companies like SK hynix, Samsung, Amazon, and AWS. Employees report a 3.8/5.0 employer rating (Work-Life 4.0, Compensation 4.2, Culture 3.5).

Application questions