Success With Microservices

The Critical C's


Barry S. Stahl

Solution Architect & Developer

@bsstahl@cognitiveinheritance.com

https://CognitiveInheritance.com

Transparent Half Width Image 720x800.png

Favorite Physicists & Mathematicians

Favorite Physicists

  1. Harold "Hal" Stahl
  2. Carl Sagan
  3. Richard Feynman
  4. Marie Curie
  5. Nikola Tesla
  6. Albert Einstein
  7. Neil Degrasse Tyson
  8. Niels Bohr
  9. Galileo Galilei
  10. Michael Faraday

Other notables: Stephen Hawking, Edwin Hubble, Leonard Susskind, Christiaan Huygens

Favorite Mathematicians

  1. Ada Lovelace
  2. Alan Turing
  3. Johannes Kepler
  4. Rene Descartes
  5. Isaac Newton
  6. Emmy Noether
  7. George Boole
  8. Blaise Pascal
  9. Johann Gauss
  10. Grace Hopper

Other notables: Daphne Koller, Grady Booch, Leonardo Fibonacci, Evelyn Berezin, Benoit Mandelbrot

Some OSS Projects I Run

  1. Liquid Victor : Media tracking and aggregation [used to assemble this presentation]
  2. Prehensile Pony-Tail : A static site generator built in c#
  3. TestHelperExtensions : A set of extension methods helpful when building unit tests
  4. Conference Scheduler : A conference schedule optimizer
  5. IntentBot : A microservices framework for creating conversational bots on top of Bot Framework
  6. LiquidNun : Library of abstractions and implementations for loosely-coupled applications
  7. Toastmasters Agenda : A c# library and website for generating agenda's for Toastmasters meetings
  8. ProtoBuf Data Mapper : A c# library for mapping and transforming ProtoBuf messages

Fediverse Supporter

Logos.png

http://GiveCamp.org

GiveCamp.png

Achievement Unlocked

bss-100-achievement-unlocked-1024x250.png

The Critical C's of Microservices

6 Conversations around building µservice architectures

Software Engineering Discussion 800x800.jpg
  Feynman 800x603.jpg

The Law of Gravitation: an Example of Physical Law by Richard Feynman from the Cornell Messenger Lectures, November 9, 1964.

Bus Maintenance System

Given: A telemetry message is sent by a bus in our network

When: The message requires maintenance action (details to be supplied)

Then: A WorkOrder is created for the bus

And: The bus manufacturer is notified

Bus-Maintenance Required-800x800.jpg
  Bus Work Order System - Monolith.png

Consistency

Data is consistent when it appears the same way when viewed from any perspective


Systems are consistent when all of the data in them is consistent

Consistency 800x800.jpg

Consistency is a Lie

All consistency is simulated


The Universe is Eventually Consistent

Lightning 800x800.jpg

Eventual Consistency

KendallMiller.png

The real world doesn't require the consistency we tend to demand of our systems - Kendall Miller 2022-01-14

  Eventual Consistency Post.png

All Models are Wrong

AllModelsAreWrong 800x800.jpg

Eventual Consistency

Eliminates temporal coupling

  • Enables Patterns & Practices that can improve

    • Resiliency & Fault Tolerance
      • Failures are isolated and self-healing
    • Maintainability & Extensibility
      • Easier to change individual components without disrupting others
    • Performance & Scalability
      • Components operate and scale independently
Eventual Consistency 800x800.jpg

Store & Forward

Get it safe, then get it done

  • Do no work until the data is persisted
  • Verify the message is well-formed
  • Do not verify the message is valid if that requires work
  • Once stored, return "Accepted"
    • i.e. HTTP 202
SafetyFirst 800x800.jpg
  Bus Work Order System - Monolith.png
  Bus Management System - Simplified - Consistency-StoreAndForward - 1684x800.png

Consistency

Embrace Eventual Consistency

  • There is no such thing as "Full Consistency"
  • Our systems RARELY need more than eventual consistency
  • Don't force your users or other systems to wait for an answer
  • Get it safe, then get it done
Consistency Flipped 800x800.jpg

Consistency Conversation

Development teams should have conversations around Consistency that are primarily focused around making certain that the system is assumed to be eventually consistency throughout.

Software Engineering Discussion 800x800.jpg

Consistency Questions

  • What patterns and tools will we use to create systems that support reliable, eventually consistent operations?
    • Example: Reliable messaging using ASB or Cosmos Change-Feed
  • How will we identify areas where stronger consistency has been wedged-in and should be removed?
    • Example: Look for polling, distributed trx, etc
  • How will we prevent future demands for strong consistency, either explicit or assumed, from weakening our systems?
    • Example: Document our commitment to Eventual Consistency in an ADR and explicitly discuss in code reviews
  • How will we identify when there are unusual or unacceptable delays in reaching consistency?
    • Example: Data Freshness SLIs/SLOs
  • How will we communicate the status of the system and any delays in reaching consistency to the stakeholders?
    • Example: Dashboards or toast notifications

Context

The boundary within which a system or sub-system operates

  • Actors
    • User
    • Other Sub-systems
    • Other Systems
  • Entities
    • Owned (Read/Write)
    • Reference (Read only)
Context 800x800.jpg

Event Storming

A process for modeling a business domain from the perspective of the business experts

Storm 1024x683.jpg
 

Bounded Context

A single business grammar

  • Every term has exactly one meaning
  • Example: Bus
    • Work Order Context: A vehicle needing diagostics or repair
      • Service Time
      • Repair History
    • Inventory Context: A depreciable asset
      • In-Service date
      • Manufacturer and Type
  • Terms are captured in Ubiquitous Language
Bounded Context 800x800 - Text Blurred.jpg

Execution Context

  • The unit of work of all services

  • The life-cycle of a single request

  • The key to service reliability

Execution Context - Silicon City 800x800.jpg

Two At-A-Time

Is is NOT possible to reliably make more than one change to system state in a single execution context

Two changes can be made to system state in a single execution context if:

  • The 1st change is idempotent
  • The 2nd change is unreliable
Two Things at Once 800x800.jpg
 

Less Reliable Changes

Should be rare, approved, documented & quickly remediated

  • Logging & Telemetry
  • Cache-aside

  • Legacy systems integration
  • Security or major incident
Changes to System State 800x800.jpg

Idempotence

We should make our services idempotent when reasonable to do so

  • Improves Reliability
    • Retry at will
  • May Hurt Maintainability & Extensibility
    • Once published or utilized, it becomes part of the contract
  • Can be abused
    • Unintentional DoS attacks
Idempotence is Golden 800x800.jpg

Context

We must define our Bounded Contexts well and defend our Execution contexts

  • Utilize Event Storming
  • Avoid Dual-Writes
  • Leverage Idempotence when possible
Context 800x800.jpg
  Bus Management System - Simplified - Consistency-StoreAndForward - 1684x800.png
  Bus Management System - Simplified - Context-Single Write - 1684x722.png
  Bus Management System - Simplified - Context-Producer - 1684x722.png

Context Conversation

Development teams should have conversations around Context that are primarily focused around the tools and techniques that they intend to use to define their Bounded Contexts and to avoid the Dual-Writes Anti-Pattern.

Software Engineering Discussion 800x800.jpg

Context Questions

  • What database technologies will we use and how can we leverage these tools to create downstream events based on changes to the database state?

  • Which of our services are currently idempotent and which ones could reasonably made so? How can we leverage our idempotent services to improve system reliability?

  • Do we have any services right now that contain business processes implemented in a less-reliable way? If so, pulling this functionality out into their own microservices might be a good starting point for decomposition.

  • What processes will we as a development team implement to track and manage the technical debt of having business processes implemented in a less-reliable way?

  • What processes will we implement to be sure that any future less-reliable implementations of business functionality are made only after strong consideration and with a plan to pay it off, appropriate documentation and prioritization by the business and product owner.

Contract

An agreement between a service provider and the service's consumers

  • Operations

  • Data Types & Structures

  • Quality of Service

    • i.e. SLAs
  • Versioning Strategy

    • i.e. compatibility
  • Behavior

    • i.e. idempotence
Contract 800x800.jpg

Upstream & Downstream Contract

Both directions need to be carefully considered

  • Upstream
    • How other systems pass info us
  • Downstream
    • How we pass info to other systems

  • Commands
    • Requests to modify data
  • Queries
    • Requests for data
Upstream and Downstream Contract 800x800.jpg

Sync vs Async

Synchronous interfaces considered harmful

  • Synchronous
    • Combines downstream effort
    • Rapid feedback (see Consistency)
  • Asynchronous
    • Improves Reliability
      • Avoids Temporal Coupling
      • Limits cascading failures
    • Maintains Team Agility
      • Single responsibility
    • Avoids Ineficiencies
      • Optimized Storage
      • Eliminates noisy-neighbors
Async Interfaces 800x800.jpg

Sending Messages Downstream

Use the best-fit approach to publication

  • Publish what you Know

    • Domain Events
    • Event Carried State Transfer
  • Claim Check

    • Minimal messages with links
    • Requires an API
  • Event Sourcing

    • When interpretation of events may change
  • Request-Reply

    • REST over Reliable Messaging
Messaging 800x800.jpg

The Canonical Model

Maintain a distinct internal model

  • Preserve team agility
  • Stability for Consumers
  • Robustness to Change

  • Only one service should write
  • Only services in THIS SUBSYSTEM should read
Canonical vs Public Contract 800x800.jpg

Contract

Once a message is defined, all stakeholders have a compatibility expectation

  • Protect your internal, isolated stream for rapid iteration
  • Use upstream contracts to bring external data to local stores
  • Avoid creating predicate queries for downstream use
Contract 800x800.jpg
  Bus Management System - Simplified - Context-Producer - 1684x722.png
  Bus Management System - Simplified - Contract-Upstream - 1684x722.png
  Bus Work Order System - Monolith.png
  Bus Management System - Simplified - Final - 1684x722.png

Contract Conversation

Development teams should have conversations around Contract that are primarily focused around creating processes that define any integration contracts for both upstream and downstream services, and serve to defend their internal data representations and implementations against any external consumers.

Software Engineering Discussion 800x800.jpg

Contract Questions

  • How will we isolate our internal data representations from those of our downstream consumers?
  • What types of compatibility guarantees are our tools and practices capable of providing?
  • What procedures should we have in place to monitor incoming and outgoing contracts for compatibility?
  • What should our procedures look like for making a change to a stream that has downstream consumers?
  • How can we leverage upstream messaging contracts to further reduce the coupling of our systems to our upstream dependencies?

Chaos

Software will fail -- in the worst possible way and at the worst possible time

  • Have low expectations of commodity infrastructure

    • Networks will segment
    • Pods, Servers and Drives will fail
    • Containers and OSs will become unstable
  • Remember The Fallacies of Distributed Computing

Chaos 800x800.jpg

Embrace the Chaos

Embrace the fact that failures will occur

  • Analytical Testing

    • FMEA
    • Virtual Chaos Engineering
  • Empirical Testing

    • Leverage Feature Flags
    • Simian Army (Chaos Monkey, etc)
  • Start by Testing in Lower Environments

Embrace the Chaos 800x800.jpg

Chaos Conversation

Development teams should have conversations around Chaos that are primarily focused around procedures for identifying and remediating possible failure points in the application.

Software Engineering Discussion 800x800.jpg

Chaos Questions

  • How can we make our systems self-healing?

  • How will we evaluate potential sources of failures in our systems before they are built?

    • How will we handle the inability to reach a dependency such as a database?
    • How will we handle duplicate messages sent from our upstream data sources?
    • How will we handle messages sent out-of-order from our upstream data sources?
  • How will we expose possible sources of failures during any pre-deployment testing?

  • How will we expose possible sources of failures in the production environment before they occur for users?

  • How will we identify errors that occur for users within production?

  • How will we prioritize changes to the system based on the results of these experiments?

Competencies

The skills and capabilities that differentiate us from our competitors

  • Focus on processes that
    • Are unique to our business
      • Provide competitive advantage
    • Change frequently
    • Are complex parts of the domain
      • Require frequent comms with domain experts
Competencies 800x800.jpg

Non-Core Domains

Avoid building in Supporting and Generic Domains

  • Supporting Domains

    • Inventory Management
    • Employee Relations
  • Generic Domains

    • Feature Flagging
    • Logging & Tracing
Core Competencies 800x800.jpg
 

Competencies Conversation

Development teams should have conversations around Competencies that are primarily focused around what systems, sub-systems, and components should be built, which should be installed off-the-shelf, and what libraries or infrastructure capabilities should be utilized.

Software Engineering Discussion 800x800.jpg

Competencies Questions

  • What are our core competencies?
  • How do we identify "build vs. buy" opportunities?
  • How do we make "build vs. buy" decisions on needed systems?
  • How do we identify cross-cutting concerns and infrastructure capabilites that can be leveraged?
  • How do we determine which libraries or infrastructure components will be utilized?
  • How do we manage the versioning of utilized components, especially in regard to security updates?
  • How do we document our decisions for later review?

Coalescence

Automate gathering operational data

  • Deployment & Verfication Testing
    • Know if a deployment is successful
  • Logging, Tracing, & Monitoring
    • Know if a system is unstable
    • Diagnose problems quickly
  • Feature Flagging
    • Turn off a feature quickly
  • Dependency Status
    • Know when required dependencies are unstable
Coalescence 800x800.jpg

Coalescence Conversation

Development teams should have conversations around Coalescence that are primarily focused around what data streams are available that provide insight into the system, and how to make them available easily to all who might need them.

Software Engineering Discussion 800x800.jpg

Questions About Coalescence

  • What is our mechanism for deployment and system verification?
  • What is our mechanism for logging/traceability within our system?
  • How will we increase the level of logging when needed?
  • How will we expose SLIs and other metrics so they are available when needed?
  • How will we identify the status of dependencies so we can understand when our systems are reacting to downstream anomalies?
  • Are there ways to perform ad-hoc queries against logs, metrics or dependencies to provide additional insight in an outage?
  • How will we know when there are anomalies in our logs, metrics or dependencies?
  • How will we coalesce our logs, metrics and dependency statuses for easy access?

Recommendations on these Conversations

  • Start having them early

    • Later today will do fine
  • Don't try to do them all at once

    • Never more than 90 min at a time
    • Usually 30-60 min
  • Get help from local experts

  • Leverage Domain Driven Design

Recommendations on the Conversations 800x800.jpg

The Critical C's of Microservices

6 Conversations around building µservice architectures

Software Engineering Discussion 800x800.jpg

Resources

SuccessWithMicroservices_QR.png

Appendix A - Uptime Metrics

  • 3-9's (unimportant)

    • 99.9% uptime
    • 1 outage second in 1000
    • ~ 43 min / month
  • 4-9's (important)

    • 99.99% uptime
    • 1 outage second in 10,000
    • ~ 4.3 min / month
  • 5-9's (critical)

    • 99.999% uptime
    • 1 outage second in 100,000
    • ~ 0.43 min (26 sec) / month

Appendix B - Delivery Guarantees

  • At-Least Once

    • Every messages will be delivered 1+ times
    • The vast majority of message systems make this guarantee
  • At-Most once

    • Every message will be delivered 0-1 times
    • i.e. Logging
  • Exactly once

    • Every message will be delivered 1 and only 1 time
    • Very difficult and expensive
    • Limited usefulness while maintaining this guarantee
    • Combination of idempotent input and a downstream transaction