Job Title
Site Reliability Engineer (SRE) – Cloud Platform
Location
Japan
Workplace Type
Hybrid / Global Team Environment
Industry
Cloud Computing / Telecommunications Technology / Enterprise Software
Job Category
Cloud Infrastructure Engineering / Site Reliability Engineering (SRE)
Career Level
Mid to Senior Level
Company Overview
A global cloud technology organization is seeking a Site Reliability Engineer (SRE) to support enterprise cloud platform operations and customer success initiatives. The company provides cloud infrastructure and platform solutions that enable organizations to deploy, manage, and optimize mission-critical workloads in secure and scalable environments.
Role Overview
This role is responsible for designing, supporting, and optimizing cloud platform environments while ensuring reliability, scalability, security, and operational excellence. Working within a global customer success organization, the successful candidate will collaborate with customers, architects, and engineering teams to support cloud adoption, infrastructure operations, and complex technical initiatives.
Department Overview
The Customer Success organization focuses on helping enterprise customers maximize the value of their cloud investments through technical guidance, operational support, and architectural expertise.
The team operates globally across multiple regions and consists of cloud architects, engineers, and program management professionals working in a highly collaborative environment.
Key Responsibilities
Cloud Architecture & Platform Reliability
Design secure, scalable, and highly available cloud architectures.
Support infrastructure modernization and cloud transformation initiatives.
Improve platform reliability, resilience, and operational efficiency.
Optimize cloud environments to meet customer and business requirements.
Identify and mitigate infrastructure risks and performance bottlenecks.
Site Reliability Engineering & Operations
Maintain and support cloud infrastructure environments.
Troubleshoot complex platform, operating system, and infrastructure issues.
Drive operational excellence through automation and process improvement.
Monitor system health, availability, and performance metrics.
Participate in incident management, root cause analysis, and service restoration activities.
Customer Success & Technical Support
Provide technical support and infrastructure expertise to enterprise customers.
Assist customers with cloud adoption, optimization, and operational best practices.
Support mission-critical cloud workloads and production environments.
Collaborate with customers and internal stakeholders to ensure successful service delivery.
Automation & Infrastructure Management
Develop automation solutions to improve operational efficiency.
Implement infrastructure configuration management and orchestration processes.
Create scripts and tools to streamline deployment, monitoring, and maintenance activities.
Support Infrastructure-as-Code and automated operational frameworks.
Security & Compliance
Support cloud security initiatives and best practices.
Implement Identity and Access Management (IAM) controls.
Assist in vulnerability assessment and remediation activities.
Promote secure infrastructure design and operational governance.
Required Qualifications
Technical Skills
Strong Linux system administration experience.
Advanced troubleshooting and problem-solving capabilities.
Hands-on experience with Kubernetes administration and container orchestration.
Certified Kubernetes Administrator (CKA) or equivalent practical experience.
Experience with Ansible or similar configuration management tools.
Strong understanding of cloud infrastructure services, including:
Compute
Storage
Networking
Automation Platforms
Knowledge of Identity and Access Management (IAM).
Experience with virtualization technologies.
Experience supporting containerized environments.
Strong scripting and automation skills using:
Bash
Python
Other infrastructure automation tools
Education
Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.
Equivalent practical experience will also be considered.
Preferred Qualifications
Experience supporting large-scale enterprise cloud environments.
Experience within telecommunications, cloud platform, or infrastructure service providers.
Knowledge of cloud-native architecture and distributed systems.
Experience supporting customer-facing cloud operations.
Bachelor's degree in Engineering, Technology, Computer Applications, or related fields.
Language Requirements
Fluent English communication skills.
Ability to collaborate effectively with global teams and international stakeholders.
Ideal Candidate Profile
Passionate about cloud technologies, automation, and infrastructure reliability.
Strong analytical and troubleshooting mindset.
Self-motivated and capable of working independently.
Quick learner with the ability to adapt to evolving technologies.
Customer-focused with strong communication and collaboration skills.
Comfortable operating within global and cross-functional environments.
Interested in driving innovation through modern cloud technologies and operational excellence.
What Makes This Opportunity Attractive
Opportunity to work on large-scale cloud infrastructure supporting enterprise customers.
Exposure to cloud-native technologies, Kubernetes, automation, and platform engineering.
Collaboration with globally distributed engineering and customer success teams.
Strong focus on innovation, scalability, reliability, and customer outcomes.
Opportunity to influence cloud architecture and operational best practices in a growing cloud platform environment.