18. Using sos report Presets.

18. Using sos report Presets.

Integration of sos report in Incident Management Pipelines.

Monitoring and observability tools — Grafana, Prometheus, traces, logs — tell you that something is wrong and where. They do not tell you what the host operating system was doing at that moment: which processes were consuming memory, what the kernel OOM killer decided, whether a filesystem was saturated, what the block device queue looked like, what firewall rules were in effect. That data lives on the node, is often ephemeral, and disappears or changes as the system recovers.

The purpose of integrating sos report into the pipeline is to capture that OS-level snapshot automatically, at the moment of the alert, before the evidence degrades without requiring a human to log into the node and collect it manually.

More specifically it achieves four things:

  • Speed of diagnosis. The data is already collected and analysed by the time the SRE opens the alert. They review findings instead of gathering evidence.

  • Evidence preservation. Memory state, kernel ring buffer entries, and process tables are ephemeral. Automated collection catches them before the system recovers and overwrites them.

  • Reduced toil. Manual OS diagnostics during an incident are slow, error-prone, and inconsistent between engineers. Presets make the collection reproducible and automatic.

  • Completeness. Every incident of the same type produces the same shape of data, making cross-incident comparison and pattern recognition — including by an AI analysis tool  meaningful and reliable.

In short: monitoring tells you the what, tracing tells you the where, and sos report presets tell you the why   automatically, consistently, and fast enough to be useful during the incident rather than after it.

In the previous article we covered how /etc/sos/sos.conf eliminates the need to repeat long command-line invocations by defining a system-wide default configuration for every sos report run. That mechanism is excellent for establishing a consistent baseline across your fleet — but it applies the same behavior to every execution.

Real-world production environments, however, rarely need the same diagnostic data every time. A disk saturation alert demands completely different information than an OOM kill event or a degraded microservice. Forcing a full-blown sos report in response to every alert is wasteful: it takes longer, generates larger archives, and buries the relevant data under a mountain of unrelated information.

This is exactly the problem that sos report presets are designed to solve and what is covered in this article..


TL;DR

sos report presets are named, reusable configurations stored in /etc/sos/presets.d/ that tell sos report exactly which plugins to run and which options to apply — triggered with a single --preset <name> flag. In a Kubernetes environment, presets are deployed to worker nodes (not pods) and executed automatically when Grafana fires an alert: Grafana calls a webhook, the webhook triggers an Ansible playbook, the playbook SSHes into the affected node and runs sos report with the matching preset, and the resulting archive is encrypted and uploaded to sos-vault for AI analysis — all within 20 – 45 seconds.

The five presets covered in this article:

Preset

Alert trigger

Key plugins

Est. time

Est. size

diskProblem

Filesystem > 85%

or I/O latency spike

filesystem, block, lvm2,

zfs, scsi, cifs, logs

20 – 45 s

3 – 8 MiB

s3Unavailable

S3 error rate spike

networking, firewalld,

nftables, dns, logs

20 – 40 s

3 – 7 MiB

databaseUnreachable

DB connection timeout

or pool exhaustion

networking, firewalld,

nftables, ulimits, pacemaker,

logs

20 – 40 s

3 – 7 MiB

OOMKilled

OOMKilled container

via kube-state-metrics

memory, cgroups, logs

15 – 35 s

2 – 6 MiB

serviceDegradation

Elevated error rate or

P99 latency

networking, cpu, systemd,

interrupts, ulimits, logs

20 – 45 s

3 – 8 MiB

Every preset also includes host, kernel, memory, process, release, and date as a fixed baseline, and caps both log-size and journal-size at 15 MB to keep archives predictable. For comparison, an unconstrained full sos report on the same node takes 3 – 8 minutes and produces 50 – 200 MiB.

The minimal JSON structure:

{
  "presetName": {
    "desc": "Short description",
    "note": "Behavioural note",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,...",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  }
}

Read on for the full pipeline architecture, all five preset definitions, the complete deployable JSON file, and the Ansible playbooks to wire everything together.


Supported distributions

The preset functionality described in this article applies to the following distributions: Red Hat Enterprise Linux, CentOS, Rocky Linux, AlmaLinux, Amazon Linux, Azure Linux (CBL-Mariner), Circle Linux, CloudLinux, Anolis OS, OpenCloudOS, OpenEuler, and UnionTech Server. Custom presets do not work out of the box on Ubuntu, Debian, SUSE, openSUSE, or Google Container-Optimized OS.

Introduction: One Command, Many Faces

A preset is a named, reusable configuration that instructs sos report to behave in a precisely defined way — limiting which plugins run, which options are applied, and how the resulting archive is handled — all without the operator ever having to type a complex command.

When paired with an incident management pipeline, presets become something far more powerful: they turn sos report into a targeted, automated diagnostic probe that can be fired in response to a specific alert, collect only the data relevant to that alert, and ship a compact, encrypted report to an analysis platform — all within seconds.

For DevOps and SRE teams, this means:

  • Reduced mean time to data (MTTD): Presets are fast because they only run the plugins you need.

  • Smaller archives: A focused report can be an order of magnitude smaller than a full collection.

  • Automation-friendly: A preset can be triggered by name with a single command, making it trivial to integrate with Ansible playbooks, Kubernetes operators, or any scripting layer.

  • Reproducible diagnostics: The same preset produces the same shape of data every time, making cross-incident comparisons meaningful.

This article walks you through everything you need to know to build, deploy, and integrate sos presets into a real Kubernetes production environment.


How Presets Work

Presets are defined in JSON files stored under the /etc/sos/presets.d/ directory. Each file can contain one or more named presets, each specifying a set of options that will be applied when sos report is invoked with the --preset <name> flag.

When sos report runs with a preset:

  1. It loads built-in defaults.

  2. It reads /etc/sos/sos.conf (if present).

  3. It applies the named preset, overriding any matching options from sos.conf.

  4. It applies any command-line flags last (highest priority).

Precedence order:

Command-line options > preset > sos.conf > built-in defaults

This layered model means you can use /etc/sos/sos.conf to define truly global defaults (upload parameters, upload and encryption keys, compression type, temporary directory) and use presets to express the intent of each diagnostic — which plugins to run, which data to upload, and under what constraints — with presets always able to override the baseline when needed.


The Preset JSON File

File Location

Preset files are placed in /etc/sos/presets.d/. Any .json file in this directory is automatically discovered and loaded when sos report starts. Presets are organized into multiple files — one per team, one per application, one per environment or as in this article one per type of alert — .

How the File Is Parsed

It is important to understand the exact structure the sos parser expects. The sos report command reads each JSON file in search of a preset name. Under each preset name, it looks for three specific keys:

JSON key

Purpose

desc

Short description shown by --list-presets

note

Behavioural note shown by --list-presets

args

A dictionary of sos option names and their values

Plugin selection is expressed as an entry inside args. Please note that the args key names use underscores in contrast of the dashes that are used in the command line and the sos.conf file — they correspond directly to the regular command line options (e.g. only-plugins, log-size, encrypt-pass, upload-url).

To easily find out what plugins you need for a specific purpose or to retrieve an specific command output, log file or configuration file; there is a plugin finder tool available in this article that allows you to find plugins by name, by included files, by included commands and describes in great detail what each plugin does and what options it supports.

Minimal Structure

{
  "preset-name": {
    "desc": "Short description of this preset",
    "note": "Behavioural note displayed with --list-presets",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  }
}

Minimum Required Fields in Every Preset

Regardless of the alert type, every preset in this guide includes:

Option key

Purpose

batch

Suppresses interactive prompts for automated execution

log-size

Caps the size of each collected log file (MB)

journal-size

Caps the size of the systemd journal collected (MB) — critical for keeping archives predictable; omitting this leaves the default at 100 MB and is the single biggest driver of archive bloat

encrypt-pass

Encrypts the archive before upload

upload-url

Destination endpoint on sos-vault

upload-user

Authentication identity

upload-pass

Authentication credential

upload-method

HTTP method used by the uploader

only-plugins (includes host, kernel, memory, process, release, date)

Core system context — always included so every report can be correlated to a specific host, kernel, and point in time

NOTE: As mentioned before options like encrypt_pass, upload_url, upload_user, upload_pass, and upload_method would be better to be defined in the /etc/sos/sos.conf file as they all are general and applicable to all presets however this article defines them in each preset JSON file.


Creating Presets

There are two ways to create presets: using the sos report command itself, or manually editing a JSON file.

Option 1: Using the sos Command

The sos report command can generate preset entries from command-line options. This is convenient for interactive, one-off preset creation, because it validates options before writing:

sudo sos report --add-preset diskProblem \
  --desc "Disk and filesystem diagnostic for storage-related alerts" \
  --note "Triggered by Grafana disk usage alert. Includes filesystem, block, lvm2, and storage plugins." \
  --batch \
  --log-size=15 \
  --journal-size=15 \
  --only-plugins=host,kernel,memory,process,release,date,filesystem,block,cifs,scsi,logs,lvm2,zfs \
  --encrypt-pass="ENCRYPTION_PASSWORD" \
  --upload-url="https://sos-vault.com/api/upload" \
  --upload-user="[email protected]" \
  --upload-pass="UPLOAD_PASSWORD" \
  --upload-method=post

This writes the preset into a JSON file under /etc/sos/presets.d/. The --desc and --note fields are visible when listing presets and serve as inline documentation for your team — always include them.

Option 2: Manual JSON Creation

You can also create the preset file directly. This is the preferred approach when deploying presets at scale via configuration management tools (Ansible, Puppet, Chef) or when building a standard preset library to distribute across a fleet of nodes. Create the file at /etc/sos/presets.d/k8s-production.json.

Listing active presets is straightforward:

sudo sos report --list-presets

A Real-World Scenario: Kubernetes + Grafana + Ansible

To demonstrate how presets deliver real operational value, let us walk through a concrete production environment.

Where sos report Runs in a Kubernetes Environment

Before describing the pipeline, it is worth being precise about one architectural point: sos report runs on the Kubernetes node, not inside a pod.

sos report requires root, is not installed in application container images, and collects data from the host operating system — the physical or virtual machine running the Kubernetes worker. Disk saturation, OOM kills, LVM volume state, block device errors, and kernel ring buffer output are all node-level phenomena, invisible from inside a container's namespaced view. The preset files under /etc/sos/presets.d/ are therefore deployed to every Kubernetes worker node, not to pods.

When Grafana fires an alert about a pod — say, an OOMKilled event — the pipeline must identify which node that pod was scheduled on and run sos report on that node. The Kubernetes API makes this trivial: every pod's status includes a nodeName field that names the worker it landed on.

The setup:

  • A Kubernetes cluster running multiple production workloads.

  • Grafana for metrics collection, dashboarding, and alert management.

  • Ansible for agentless SSH-based remote command execution across cluster nodes.

  • sos-vault 2.0.0 as the central repository for diagnostic archives, with its built-in AI analysis engine to automatically parse incoming reports and surface findings.

The Pipeline

sos report pipeline flow integration

Step 1: Deploy Presets to Every Worker Node

Use Ansible itself to distribute the preset file to every node in the cluster, ensuring it is in place before any alert fires:

- name: Deploy sos report presets to all Kubernetes nodes
  hosts: k8s_workers
  become: true
  tasks:
    - name: Ensure presets directory exists
      file:
        path: /etc/sos/presets.d
        state: directory
        mode: '0755'

    - name: Deploy k8s-production preset file
      copy:
        src: files/k8s-production.json
        dest: /etc/sos/presets.d/k8s-production.json
        owner: root
        group: root
        mode: '0600'

Step 2: Write a Playbook Per Preset

Each alert type maps to a short playbook. The playbook receives the affected node's hostname and case ID as variables — both supplied by the Grafana webhook:

# playbooks/sos-disk-problem.yml
- name: Run diskProblem sos report on affected node
  hosts: "{{ target_node }}"
  become: true
  tasks:
    - name: Execute sos report with diskProblem preset
      command: >
        sos report
        --preset diskProblem
        --case-id="{{ case_id }}"
      async: 300
      poll: 10

Invoke it from the webhook handler:

ansible-playbook playbooks/sos-disk-problem.yml \
  -e "target_node=worker-node-03 case_id=INC-20240517-001"

Step 3: Configure Grafana to Fire the Webhook

In Grafana, configure a contact point of type Webhook. Point it at a lightweight webhook receiver (a small Flask or FastAPI service, or a serverless function) that extracts the alert labels, queries the Kubernetes API for the pod's nodeName, and invokes the correct Ansible playbook.

A minimal webhook handler in Python illustrates the logic:

from flask import Flask, request
import subprocess, json

app = Flask(__name__)

PRESET_MAP = {
    "DiskUsageHigh":         "sos-disk-problem.yml",
    "S3ErrorRateHigh":       "sos-s3-unavailable.yml",
    "DBConnectionTimeout":   "sos-database-unreachable.yml",
    "OOMKilled":             "sos-oom-killed.yml",
    "ServiceLatencyHigh":    "sos-service-degradation.yml",
}

@app.route("/grafana-alert", methods=["POST"])
def handle_alert():
    payload  = request.json
    alert    = payload["alerts"][0]
    rule     = alert["labels"]["alertname"]
    pod      = alert["labels"].get("pod", "")
    case_id  = alert["fingerprint"]

    # Resolve the node name from the Kubernetes API
    node = get_node_for_pod(pod)
    playbook = PRESET_MAP.get(rule)

    if playbook and node:
        subprocess.Popen([
            "ansible-playbook", f"playbooks/{playbook}",
            "-e", f"target_node={node} case_id={case_id}"
        ])

    return "", 200

Step 4: Let sos-vault Do the Rest

Once sos report completes on the node, the archive is encrypted, uploaded, and waiting in sos-vault. The sos-vault 2.0.0 AI analysis engine immediately begins processing the report, running automated checks across all collected data: it correlates disk usage metrics with LVM volume group state, checks for filesystem errors in the kernel ring buffer, identifies which processes are the top consumers of the saturated path, and surfaces findings in a structured incident view.

By the time the on-call SRE opens the alert, the diagnostic cycle is already done. The SRE does not collect data — they review data. That shift is the real value of this pipeline.


The Five Presets

The following presets cover the five most common alert categories in this environment. All five are deployed to every worker node via the Ansible playbook above.

Preset 1: diskProblem

Alert trigger: Grafana detects filesystem usage above 85% or a sharp spike in disk I/O latency.

Goal: Collect everything needed to understand the storage layer on the affected node — which filesystems are full, which logical volumes are involved, which block devices are saturated, and what the system logs say about it.

Plugins included: host, kernel, memory, process, release, date, filesystem, block, cifs, scsi, logs, lvm2, zfs

Creating the preset:

sudo sos report --add-preset diskProblem \
  --desc "Disk and filesystem diagnostic for storage saturation alerts" \
  --note "Triggered by Grafana: filesystem_usage > 85% or disk I/O latency spike. Covers block, lvm2, zfs, cifs, scsi, and filesystem layers." \
  --batch \
  --log-size=15 \
  --journal-size=15 \
  --only-plugins=host,kernel,memory,process,release,date,filesystem,block,cifs,scsi,logs,lvm2,zfs \
  --encrypt-pass="ENCRYPTION_PASSWORD" \
  --upload-url="https://sos-vault.com/api/upload" \
  --upload-user="[email protected]" \
  --upload-pass="UPLOAD_PASSWORD" \
  --upload-method=post

JSON definition:

{
  "diskProblem": {
    "desc": "Disk and filesystem diagnostic for storage saturation alerts",
    "note": "Triggered by Grafana: filesystem_usage > 85% or disk I/O latency spike.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,filesystem,block,cifs,scsi,logs,lvm2,zfs",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  }
}

Ansible invocation:

ansible-playbook playbooks/sos-disk-problem.yml \
  -e "target_node=worker-node-03 case_id=INC-20240517-001"

Preset 2: s3Unavailable

Alert trigger: Grafana detects S3 API errors above a threshold — connection timeouts, HTTP 5xx responses, or bucket access failures from application pods.

Goal: Understand the network path between the node and the S3 endpoint. Capture DNS resolution behavior, firewall rules, active connections, and any routing anomalies at the node level.

Plugins included: host, kernel, memory, process, release, date, networking, firewalld, nftables, dns, logs

Creating the preset:

sudo sos report --add-preset s3Unavailable \
  --desc "Network and DNS diagnostic for S3 connectivity failures" \
  --note "Triggered by Grafana: S3 error rate spike. Covers networking stack, DNS resolution, firewall rules, and active connections." \
  --batch \
  --log-size=15 \
  --journal-size=15 \
  --only-plugins=host,kernel,memory,process,release,date,networking,firewalld,nftables,dns,logs \
  --encrypt-pass="ENCRYPTION_PASSWORD" \
  --upload-url="https://sos-vault.com/api/upload" \
  --upload-user="[email protected]" \
  --upload-pass="UPLOAD_PASSWORD" \
  --upload-method=post

JSON definition:

{
  "s3Unavailable": {
    "desc": "Network and DNS diagnostic for S3 connectivity failures",
    "note": "Triggered by Grafana: S3 error rate spike.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,networking,firewalld,nftables,dns,logs",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  }
}

Ansible invocation:

ansible-playbook playbooks/sos-s3-unavailable.yml \
  -e "target_node=worker-node-07 case_id=INC-20240517-002"

Preset 3: databaseUnreachable

Alert trigger: Grafana detects database connection pool exhaustion, query timeouts exceeding SLA thresholds, or complete loss of connectivity to the database endpoint.

Goal: Capture the full picture around database reachability at the OS level — network routes, socket state, and resource limits. This preset deliberately avoids database-internal plugins (which would require credentials and could be slow) and focuses on what the node itself can tell you about why the database cannot be reached.

Plugins included: host, kernel, memory, process, release, date, networking, firewalld, nftables, ulimits, logs, pacemaker

Creating the preset:

sudo sos report --add-preset databaseUnreachable \
  --desc "OS-level diagnostic for database connectivity failures" \
  --note "Triggered by Grafana: DB connection timeout or pool exhaustion. Focuses on network, socket state, resource limits, and system logs." \
  --batch \
  --log-size=15 \
  --journal-size=15 \
  --only-plugins=host,kernel,memory,process,release,date,networking,firewalld,nftables,ulimits,logs,pacemaker \
  --encrypt-pass="ENCRYPTION_PASSWORD" \
  --upload-url="https://sos-vault.com/api/upload" \
  --upload-user="[email protected]" \
  --upload-pass="UPLOAD_PASSWORD" \
  --upload-method=post

JSON definition:

{
  "databaseUnreachable": {
    "desc": "OS-level diagnostic for database connectivity failures",
    "note": "Triggered by Grafana: DB connection timeout or pool exhaustion.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,networking,firewalld,nftables,ulimits,logs,pacemaker",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  }
}

Ansible invocation:

ansible-playbook playbooks/sos-database-unreachable.yml \
  -e "target_node=worker-node-02 case_id=INC-20240517-003"

Preset 4: OOMKilled

Alert trigger: Grafana detects an OOMKilled event via kube-state-metrics (kube_pod_container_status_last_terminated_reason == "OOMKilled").

Goal: This is a node-level memory pressure event. The data needed is precise: current memory allocation on the node, cgroup memory limits, kernel OOM kill logs, and page reclaim activity. Speed matters here because the memory state is ephemeral — you want the report collected on the node as close to the event as possible.

Plugins included: host, kernel, memory, process, release, date, cgroups, logs

Creating the preset:

sudo sos report --add-preset OOMKilled \
  --desc "Memory pressure and OOM kill diagnostic" \
  --note "Triggered by Grafana: OOMKilled container detected via kube-state-metrics. Captures node-level cgroup limits, memory state, and OOM kernel logs." \
  --batch \
  --log-size=15 \
  --journal-size=15 \
  --only-plugins=host,kernel,memory,process,release,date,cgroups,logs \
  --encrypt-pass="ENCRYPTION_PASSWORD" \
  --upload-url="https://sos-vault.com/api/upload" \
  --upload-user="[email protected]" \
  --upload-pass="UPLOAD_PASSWORD" \
  --upload-method=post

JSON definition:

{
  "OOMKilled": {
    "desc": "Memory pressure and OOM kill diagnostic",
    "note": "Triggered by Grafana: OOMKilled container. Captures node-level cgroup limits, memory state, and OOM kernel logs.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,cgroups,logs",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  }
}

Ansible invocation:

ansible-playbook playbooks/sos-oom-killed.yml \
  -e "target_node=worker-node-05 case_id=INC-20240517-004"

Preset 5: serviceDegradation

Alert trigger: Grafana detects elevated error rates, increased P99 latency, or reduced throughput on a service endpoint — without a clear single root cause.

Goal: Cast a wider net at the node level. Capture CPU scheduling, interrupt pressure, network throughput, and service logs. This preset runs more plugins than the others but remains scoped — no storage collection, no database layers.

Plugins included: host, kernel, memory, process, release, date, networking, cpu, systemd, logs, interrupts, ulimits

Creating the preset:

sudo sos report --add-preset serviceDegradation \
  --desc "Broad system diagnostic for unexplained service degradation" \
  --note "Triggered by Grafana: elevated error rate or latency spike without clear root cause. Covers CPU, networking, systemd, and interrupt pressure." \
  --batch \
  --log-size=15 \
  --journal-size=15 \
  --only-plugins=host,kernel,memory,process,release,date,networking,cpu,systemd,logs,interrupts,ulimits \
  --encrypt-pass="ENCRYPTION_PASSWORD" \
  --upload-url="https://sos-vault.com/api/upload" \
  --upload-user="[email protected]" \
  --upload-pass="UPLOAD_PASSWORD" \
  --upload-method=post

JSON definition:

{
  "serviceDegradation": {
    "desc": "Broad system diagnostic for unexplained service degradation",
    "note": "Triggered by Grafana: elevated error rate or P99 latency spike.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,networking,cpu,systemd,logs,interrupts,ulimits",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  }
}

Ansible invocation:

ansible-playbook playbooks/sos-service-degradation.yml \
  -e "target_node=worker-node-01 case_id=INC-20240517-005"

Estimated Execution Time and Archive Size

One of the most common questions about integrating sos report into an incident pipeline is: how much overhead does this actually add? The table below provides realistic estimates grounded in measured data from a real system.

Methodology

The estimates below are derived from benchmarks run on an 8-core, 16 GB RAM workstation at 1% CPU load using sos 4.9.0 on Ubuntu 22.04. The key findings from that dataset are:

  • A full sos report (81 plugins, log-size=10, journal-size=10) completed in ~31 seconds and produced an ~9 MiB encrypted archive.

  • A focused 9-plugin run with the same log limits completed in ~15 seconds and produced a ~2.8 MiB archive.

  • The dominant insight: execution time is driven by a fixed startup/teardown/compression overhead of roughly 15 seconds, not by plugin count. Going from 9 to 81 plugins only adds another 15 seconds. Plugin count is therefore a minor factor in time but a more significant factor in archive size.

  • The journal is the biggest wildcard for size. Without an explicit journal_size cap, the default is 100 MiB. A busy production node generating verbose logs can push archive size from 3 MiB to 50+ MiB in seconds. This is why every preset in this guide explicitly sets journal_size: 15.

Production Kubernetes worker nodes will be slower and produce larger archives than the benchmark workstation: they carry heavier logs, more running processes, larger filesystem tables, and busier network state. The estimates below apply a conservative multiplier to account for this.

Estimates Per Preset

Preset

Plugins

Est. execution time

Est. archive size

diskProblem

13

20 – 45 s

3 – 8 MiB

s3Unavailable

11

20 – 40 s

3 – 7 MiB

databaseUnreachable

12

20 – 40 s

3 – 7 MiB

OOMKilled

8

15 – 35 s

2 – 6 MiB

serviceDegradation

12

20 – 45 s

3 – 8 MiB

For comparison: an unconstrained full sos report on the same production node would typically take 3 – 8 minutes and produce a 50 – 200 MiB archive.

What Pushes Numbers Higher

The upper end of the ranges above applies when: the node has been running for weeks without a log rotation, the logs plugin finds many large files in /var/log, systemd has a dense journal, or the filesystem plugin is iterating a large number of mount points. The diskProblem preset is the most likely to hit the upper end of both time and size, precisely because it includes the filesystem and block device layers that are active during a disk saturation event.

What Keeps Numbers Lower

A recently recycled node, a quiet log directory, or a node with a read-only root filesystem (common in some Kubernetes distributions) will produce reports toward the lower end. The OOMKilled preset is the leanest by design: it targets memory state that must be captured immediately, so it deliberately includes the minimum number of plugins.


All preset in a Single Preset File

It is also possible to group several preset in a single JSON file, deployed to every worker node in the cluster via the Ansible distribution playbook shown earlier:

{
  "diskProblem": {
    "desc": "Disk and filesystem diagnostic for storage saturation alerts",
    "note": "Triggered by Grafana: filesystem_usage > 85% or disk I/O latency spike.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,filesystem,block,cifs,scsi,logs,lvm2,zfs",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  },
  "s3Unavailable": {
    "desc": "Network and DNS diagnostic for S3 connectivity failures",
    "note": "Triggered by Grafana: S3 error rate spike.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,networking,firewalld,nftables,dns,logs",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  },
  "databaseUnreachable": {
    "desc": "OS-level diagnostic for database connectivity failures",
    "note": "Triggered by Grafana: DB connection timeout or pool exhaustion.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,networking,firewalld,nftables,ulimits,logs,pacemaker",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  },
  "OOMKilled": {
    "desc": "Memory pressure and OOM kill diagnostic",
    "note": "Triggered by Grafana: OOMKilled container detected via kube-state-metrics.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,cgroups,logs",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  },
  "serviceDegradation": {
    "desc": "Broad system diagnostic for unexplained service degradation",
    "note": "Triggered by Grafana: elevated error rate or P99 latency spike.",
    "args": {
      "batch": true,
      "log-size": 15,
      "journal-size": 15,
      "only-plugins": "host,kernel,memory,process,release,date,networking,cpu,systemd,logs,interrupts,ulimits",
      "encrypt-pass": "ENCRYPTION_PASSWORD",
      "upload-url": "https://sos-vault.com/api/upload",
      "upload-user": "[email protected]",
      "upload-pass": "UPLOAD_PASSWORD",
      "upload-method": "post"
    }
  }
}

Practical Considerations

Security: Preset files contain encryption passwords and upload credentials. Restrict access on every node with:

chmod 600 /etc/sos/presets.d/k8s-production.json
chown root:root /etc/sos/presets.d/k8s-production.json

Manage the actual secret values via your secrets management solution (Vault, AWS Secrets Manager, Kubernetes Secrets). The JSON file distributed by Ansible should contain placeholders that are substituted at deploy time using Ansible Vault or a lookup plugin.

Plugin availability: Not every plugin exists on every distribution or kernel version. Run sudo sos report --list-plugins on a representative node to verify that all plugins referenced in your presets are available before distributing the preset file.

Preset versioning: Treat preset files as code. Store them in version control alongside the Ansible playbooks, review changes in pull requests, and tag releases. A preset change is a change to your incident response behavior — it deserves the same rigor as a change to an alert rule.

Node SSH access: sos report requires SSH access to the Kubernetes worker nodes. In managed Kubernetes services (GKE, EKS, AKS), node SSH is possible but may require additional configuration such as OS login, bastion hosts, or AWS Systems Manager Session Manager as a transport. Confirm your node access method before integrating this pipeline.

Iterating on presets: After each incident, review what data the preset collected and what was missing. Update the plugin list accordingly. Presets should evolve with your understanding of your system.


Summary

Presets transform sos report from a general-purpose diagnostic tool into a precision instrument.

By defining a small library of focused, named configurations — one per alert category — and deploying them uniformly across every Kubernetes worker node, you gain:

  • Speed: Only the relevant plugins run. Focused reports complete in 20 – 45 seconds on a production node — compared to 3 – 8 minutes for a full unconstrained collection.

  • Size: Archives stay in the 2 – 8 MiB range per incident — compared to 50 – 200 MiB for a full report. Smaller archives mean faster uploads, lower storage costs, and faster AI analysis turnaround.

  • Automation: A single --preset <name> flag is all an Ansible playbook needs to produce a complete, encrypted, uploaded diagnostic from the affected node.

  • Consistency: Every diskProblem report has the same shape. Every OOMKilled report contains the same data. Cross-incident comparison becomes meaningful.

  • AI readiness: When reports are structured and scoped, the sos-vault 2.0.0 AI analysis engine can parse them faster and produce higher-confidence findings — because there is no noise to filter through.

Presets, combined with /etc/sos/sos.conf for shared defaults, give you a complete configuration strategy for sos report: a system-wide baseline that handles the boilerplate, and a per-alert library that handles the intent.

The result is a diagnostic pipeline that works while your team sleeps — collecting the right data, from the right node, at the right moment, and delivering it to sos-vault before the on-call engineer has finished reading the alert notification.