1
/
of
2
Muhammed Yunas Chirayath Meerankunju
Cloud-Native Reliability Engineering: Designing Scalable and Observable Kubernetes Platforms - Applying SRE Principles to OpenShift, Kubernetes, Observability, and Platform Operations
Cloud-Native Reliability Engineering: Designing Scalable and Observable Kubernetes Platforms - Applying SRE Principles to OpenShift, Kubernetes, Observability, and Platform Operations
Regular price
Rs. 1,420.00
Regular price
Sale price
Rs. 1,420.00
Quantity
Couldn't load pickup availability
Cloud-Native Reliability Engineering is a practitioner-focused guide to designing, operating, observing, and continuously improving reliable Kubernetes and OpenShift platforms. It addresses the critical gap between infrastructure that appears healthy and systems that actually deliver dependable outcomes for users. Grounded in production experience, the book connects Site Reliability Engineering principles with the operational realities of modern cloud-native environments.
It will enable you to:
- Define meaningful SLIs, SLOs, and error budgets based on real user experience.
- Understand Kubernetes and OpenShift architecture from a reliability-engineering perspective.
- Identify control-plane, Pod, networking, storage, and resource-management failure modes.
- Build effective observability across metrics, logs, traces, and Kubernetes events.
- Design actionable alerting strategies that reveal real service degradation rather than dashboard noise.
- Apply HPA, KEDA, and VPA concepts to autoscaling, capacity management, and cost optimization.
- Improve Kubernetes incident detection, diagnosis, escalation, recovery, and postmortem practices.
- Strengthen reliability for data platforms using technologies such as Kafka, Spark, and Airflow.
- Combine security and resilience through SCCs, mTLS, Vault, RBAC, and zero-trust principles.
- Use chaos engineering, reliability frameworks, operational measurement, and automation to move from reactive operations toward proactive and increasingly autonomous reliability.
Who should read?
- Site Reliability Engineers working with Kubernetes or OpenShift.
- Platform engineers responsible for enterprise container platforms.
- DevOps engineers supporting distributed cloud-native workloads.
- Kubernetes and OpenShift administrators seeking deeper reliability knowledge.
- Observability engineers working with Prometheus, OpenTelemetry, metrics, logs, and traces.
- Engineering leaders building SRE practices, governance, and reliability culture.
- Practitioners responsible for production incident management, scaling, platform resilience, or operational automation.
Share

About the Author