Improving a Distributed System Post-Incident

Add to wishlistAdded to wishlistRemoved from wishlist 0

Add to compare+

Duration	33m
level	Intermediate
Course Creator	Gremlin
Last Updated	14-Dec-22

Pluralsight

Category: DevOps

In this session, we will dive into a case study of how a team can recover and improve a distributed system after a major incident.

Add your review

Description
Reviews (0)

In this session, we will dive into a case study of how a team can recover and improve a distributed system after a major incident. Distributed systems are more prone to failure than other systems due to their incredible complexity and scale, and incidents are a fact of life with these systems. This year, my team faced a week long incident for our IP address management system which impacted out customers. From this incident, we had had to reevaluate our system’s performance & overhaul several keys areas of our codebase, as well as improve our monitoring, testing processes, database interactions, and reliability. Viewers will learn about these improvements and how they can apply them to their own systems to achieve greater reliability and performance. Additionally, viewers will learn how to effectively leverage monitoring practices to uncover inefficiencies in their system, tips for creating a testing process to properly stress your system before deploying to production, and how to rally a team together during a high-pressure incident.
Author Name: Gremlin
Author Description:
Gremlin is a Chaos Engineering service on a mission to help build a more reliable internet. Their solutions turn failure into resilience by offering engineers a fully hosted SaaS platform to safely experiment on complex systems, in order to identify weaknesses before they impact customers and cause revenue loss. Founded by CEO Kolton Andrus and CTO Matthew Fornaciari in 2016, the company has since raised $26.8Million in funding from Redpoint Ventures, Index Ventures, and Amplify Partners. Existi… more

Improving a Distributed System Post-Incident
33mins

User Reviews

0.0 out of 5

★★★★★

Write a review

There are no reviews yet.

Be the first to review “Improving a Distributed System Post-Incident” Cancel reply

Improving a Distributed System Post-Incident

Description
Reviews (0)

Start Course

All Categories

Improving a Distributed System Post-Incident

Table of Contents

User Reviews

Be the first to review “Improving a Distributed System Post-Incident” Cancel reply

COURSE PROVIDERS

CATEGORIES

Quick Links

Contact Us

Compare items

All Categories

Improving a Distributed System Post-Incident

Table of Contents

User Reviews

Be the first to review “Improving a Distributed System Post-Incident” Cancel reply

Related Products

Using GitHub Copilot with Python

Introduction to GitHub

Monitoring and Alerting AWS on CloudWatch

Diploma in Amazon Web Services 2019

DevOps Integration with Jenkins Pipelines

Authenticate and authorize user identities on GitHub

COURSE PROVIDERS

CATEGORIES

Quick Links

Contact Us

Compare items