klotz: service topology*

0 bookmark(s) - Sort by: Date ↓ / Title / - Bookmarks from other users for this tag

  1. Santosh Balaranganathan and colleagues at Atlassian describe their automated root cause analysis system that treats incident diagnosis as a correlation problem across three dimensions: signal type (metrics, logs, traces), time, and service topology. The pipeline scopes the search using OpenTelemetry-derived dependency graphs, detects anomalies independently per signal, temporally aligns co-occurring anomalies into bundles, traverses the graph to determine causal direction, and emits ranked hypotheses with human-readable narratives so responders can validate and act quickly.

    - Sequence fingerprinting collapses repeated fault patterns (the same upstream timeout replaying every few seconds) into a single bundle with a replay count, preventing dozens of identical hypotheses from obscuring the signal.
    - The team found statistical methods (MAD, percentile bands) work well enough for metrics anomaly detection and are far easier to debug than ML models; they reserve ML for log clustering and trace structural analysis.
    - The system is being extended with LLM-based orchestration to make RCA iterative—an agent can request additional telemetry, refine hypotheses, and adapt its investigation strategy across multiple steps rather than running one-shot.
    - A shared incident context anchors all signals, hypotheses, and actions per incident, feeding both a faulty-service pager that pages the right team early and an LLM-powered copilot that recommends mitigations (rollbacks, feature flag disablement) grounded in the actual diagnosis.
  2. Netflix uses an internal system called Service Topology to maintain a live, queryable dependency graph for thousands of microservices. The platform merges three distinct data sources—eBPF network flow logs (for kernel-level visibility), IPC metrics from instrumented services (for application context), and aggregated distributed traces (for request paths)—to provide engineers with a unified view of runtime connections. This architecture helps teams quickly identify the blast radius of failures, understand upstream dependencies, and resolve incidents more efficiently by visualizing how various components interact in real-time.

    - Data integration from eBPF logs, IPC metrics, and distributed traces ensures comprehensive coverage even for uninstrumented services.
    - A three-stage aggregation pipeline resolves multi-hop paths into direct application-to-application edges to simplify troubleshooting.
    - The processing architecture leverages Apache Pekko Streams across multi-region Kafka consumers.
    - The system supports sub-second response times and provides historical time-window aggregations for incident correlation.

Top of the page

First / Previous / Next / Last / Page 1 of 0 SemanticScuttle - klotz.me: Tags: service topology

About - Propulsed by SemanticScuttle