Aggregate

Aggregate metrics passing through a topology

status: stable egress: stream state: stateful
input: metrics
output: metrics
Aggregates multiple metric events into a single metric event based on a defined interval window. This helps to reduce metric volume at the cost of granularity.

Configuration

Example configurations

{
  "transforms": {
    "my_transform_id": {
      "type": "aggregate",
      "inputs": [
        "my-source-or-transform-id"
      ]
    }
  }
}
[transforms.my_transform_id]
type = "aggregate"
inputs = ["my-source-or-transform-id"]
transforms:
  my_transform_id:
    type: aggregate
    inputs:
      - my-source-or-transform-id
{
  "transforms": {
    "my_transform_id": {
      "type": "aggregate",
      "inputs": [
        "my-source-or-transform-id"
      ],
      "interval_ms": 10000,
      "measure_cpu_usage": false,
      "mode": "Auto"
    }
  }
}
[transforms.my_transform_id]
type = "aggregate"
inputs = ["my-source-or-transform-id"]
interval_ms = 10000
measure_cpu_usage = false
mode = "Auto"
transforms:
  my_transform_id:
    type: aggregate
    inputs:
      - my-source-or-transform-id
    interval_ms: 10000
    measure_cpu_usage: false
    mode: Auto

event_time

optional object

Event-time aggregation settings.

When present, metrics are grouped into buckets based on their timestamps rather than when they are processed. Omit this block to keep the default system-time behavior.

Grace period for late-arriving events, in milliseconds.

Each bucket accepts events until the system clock reaches bucket_end + allowed_lateness_ms, where bucket_end is the exclusive end of the event-time window. That cutoff is enforced when events are recorded, not only when a periodic flush runs. Once a bucket is emitted it is closed permanently; any later events whose timestamp falls inside it are dropped and counted via component_discarded_events_total.

Set to 0 for strict ordering (no late events allowed).

Examples
0
5000
30000

Maximum allowed time drift for future events, in milliseconds.

Acts as a clock-skew guard: events whose timestamp is further in the future than this many milliseconds (relative to the current system time) are dropped and counted via component_discarded_events_total. Defaults to 10 seconds.

Set to 0 to allow events at any future time.

Examples
0
60000
300000
default: 10000

event_time.missing_timestamp

optional string literal enum

How to handle events with missing timestamps.

Metrics that pass through unchanged for the configured mode do not require a timestamp. For metrics that would be bucketed:

  • drop (default) discards the event and increments component_discarded_events_total
  • use_system_time synthesizes a timestamp from the current system clock
Enum options
OptionDescription
dropDrop the event and count it via component_discarded_events_total.
use_system_timeUse the current system time as the event timestamp.
default: drop

graph

optional object

Extra graph configuration

Configure output for component when generated with graph command

graph.edge_attributes

optional object

Edge attributes to add to the edges linked to this component’s node in resulting graph

They are added to the edge as provided

A collection of graph edge attributes in graphviz DOT language, related to a single input component.
graph.edge_attributes.*.*
required string literal
A single graph edge attribute in graphviz DOT language.
Examples
{
  "color": "red",
  "label": "Example Edge",
  "width": "5.0"
}
Examples
{
  "example_input": {
    "color": "red",
    "label": "Example Edge",
    "width": "5.0"
  }
}

graph.node_attributes

optional object

Node attributes to add to this component’s node in resulting graph

They are added to the node as provided

graph.node_attributes.*
required string literal
A single graph node attribute in graphviz DOT language.
Examples
{
  "color": "red",
  "name": "Example Node",
  "width": "5.0"
}

inputs

required [string]

A list of upstream source or transform IDs.

Wildcards (*) are supported.

See configuration for more info.

Array string literal
Examples
[
  "my-source-or-transform-id",
  "prefix-*"
]

interval_ms

optional uint

The interval between flushes, in milliseconds.

Must be greater than zero. During this time frame, metrics (beta) with the same series data (name, namespace, tags, and so on) are aggregated.

default: 10000

measure_cpu_usage

optional bool

Enable CPU usage metrics for this transform.

When set to true, each poll of the transform task is timed using the OS thread CPU clock and the accumulated nanoseconds are reported as the component_cpu_usage_ns_total counter, tagged with component_id, component_kind, and component_type.

Defaults to false. Enable only for transforms where CPU attribution is needed, as it adds a clock_gettime call on every future poll.

default: false

mode

optional string literal enum

Function to use for aggregation.

Some of the functions may only function on incremental and some only on absolute metrics.

Enum options string literal
OptionDescription
AutoDefault mode. Sums incremental metrics and uses the latest value for absolute metrics.
CountCounts metrics for incremental and absolute metrics
DiffReturns difference between latest value for absolute; incremental metrics pass through unchanged.
LatestReturns the latest value for absolute metrics; incremental metrics pass through unchanged.
MaxMax value of absolute metric; incremental metrics pass through unchanged.
MeanMean value of absolute metric; incremental metrics pass through unchanged.
MinMin value of absolute metric; incremental metrics pass through unchanged.
StdevStdev value of absolute metric; incremental metrics pass through unchanged.
SumSums incremental metrics; absolute metrics pass through unchanged.
default: Auto

Input Types

The following table lists all telemetry data types supported by the component across possible configurations. Be aware that the available data types may differ based on the specified codec configuration.

Metrics

The following metrics are supported:
counter distribution gauge histogram set summary

Outputs

<component_id>

Default output stream of the component. Use this component’s ID as an input to downstream transforms and sinks.

Output Types

Metrics

The modified input metric event.

Telemetry

Metrics

link

aggregate_events_recorded_total

counter
The number of events recorded by the aggregate transform.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

aggregate_failed_updates_total

counter
The number of failed metric updates, incremental adds, encountered by the aggregate transform.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

aggregate_flushes_total

counter
The number of flushes done by the aggregate transform.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

component_discarded_events_total

counter
The number of events dropped by this component.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
intentional
True if the events were discarded intentionally, like a filter transform, or false if due to an error.
pid optional
The process ID of the Vector instance.

component_errors_total

counter
The total number of errors encountered by this component.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
error_type
The type of the error
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.
stage
The stage within the component at which the error occurred.

component_latency_mean_seconds

gauge

The mean elapsed time, in fractional seconds, that an event spends in a single transform.

This includes both the time spent queued in the transform’s input buffer and the time spent executing the transform itself.

This value is smoothed over time using an exponentially weighted moving average (EWMA).

host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

component_latency_seconds

histogram

The elapsed time, in fractional seconds, that an event spends in a single transform.

This includes both the time spent queued in the transform’s input buffer and the time spent executing the transform itself.

host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

component_received_event_bytes_total

counter
The number of event bytes accepted by this component either from tagged origins like file and uri, or cumulatively from other origins.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
container_name optional
The name of the container from which the data originated.
file optional
The file from which the data originated.
host optional
The hostname of the system Vector is running on.
mode optional
The connection mode used by the component.
peer_addr optional
The IP from which the data originated.
peer_path optional
The pathname from which the data originated.
pid optional
The process ID of the Vector instance.
pod_name optional
The name of the pod from which the data originated.
uri optional
The sanitized URI from which the data originated.

component_received_events_count

histogram

A histogram of the number of events passed in each internal batch in Vector’s internal topology.

Note that this is separate than sink-level batching. It is mostly useful for low level debugging performance issues in Vector due to small internal batches.

component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
container_name optional
The name of the container from which the data originated.
file optional
The file from which the data originated.
host optional
The hostname of the system Vector is running on.
mode optional
The connection mode used by the component.
peer_addr optional
The IP from which the data originated.
peer_path optional
The pathname from which the data originated.
pid optional
The process ID of the Vector instance.
pod_name optional
The name of the pod from which the data originated.
uri optional
The sanitized URI from which the data originated.

component_received_events_total

counter
The number of events accepted by this component either from tagged origins like file and uri, or cumulatively from other origins.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
container_name optional
The name of the container from which the data originated.
file optional
The file from which the data originated.
host optional
The hostname of the system Vector is running on.
mode optional
The connection mode used by the component.
peer_addr optional
The IP from which the data originated.
peer_path optional
The pathname from which the data originated.
pid optional
The process ID of the Vector instance.
pod_name optional
The name of the pod from which the data originated.
uri optional
The sanitized URI from which the data originated.

component_sent_event_bytes_total

counter
The total number of event bytes emitted by this component.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
output optional
The specific output of the component.
pid optional
The process ID of the Vector instance.

component_sent_events_total

counter
The total number of events emitted by this component.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
output optional
The specific output of the component.
pid optional
The process ID of the Vector instance.

transform_buffer_max_byte_size

gauge
The maximum number of bytes the buffer that feeds into a transform can hold.
Deprecated
This metric has been deprecated in favor of transform_buffer_max_size_bytes.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

transform_buffer_max_event_size

gauge
The maximum number of events the buffer that feeds into a transform can hold.
Deprecated
This metric has been deprecated in favor of transform_buffer_max_size_events.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

transform_buffer_max_size_bytes

gauge
The maximum number of bytes the buffer that feeds into a transform can hold.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

transform_buffer_max_size_events

gauge
The maximum number of events the buffer that feeds into a transform can hold.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

transform_buffer_utilization

histogram
The utilization level of the buffer that feeds into a transform.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

transform_buffer_utilization_level

gauge
The current utilization level of the buffer that feeds into a transform.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

transform_buffer_utilization_mean

gauge
The mean utilization level of the buffer that feeds into a transform. This value is smoothed over time using an exponentially weighted moving average (EWMA).
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

utilization

gauge
A ratio from 0 to 1 of the load on a component. A value of 0 would indicate a completely idle component that is simply waiting for input. A value of 1 would indicate a that is never idle. This value is updated every 5 seconds.
component_id
The Vector component ID.
component_kind
The Vector component kind.
component_type
The Vector component type.
host optional
The hostname of the system Vector is running on.
pid optional
The process ID of the Vector instance.

Examples

Aggregate over 5 seconds

Given this event...
[{"metric":{"counter":{"value":1.1},"kind":"incremental","name":"counter.1","tags":{"host":"my.host.com"},"timestamp":"2021-07-12T07:58:44.223543Z"}},{"metric":{"counter":{"value":2.2},"kind":"incremental","name":"counter.1","tags":{"host":"my.host.com"},"timestamp":"2021-07-12T07:58:45.223543Z"}},{"metric":{"counter":{"value":1.1},"kind":"incremental","name":"counter.1","tags":{"host":"different.host.com"},"timestamp":"2021-07-12T07:58:45.223543Z"}},{"metric":{"counter":{"value":22.33},"kind":"absolute","name":"gauge.1","tags":{"host":"my.host.com"},"timestamp":"2021-07-12T07:58:47.223543Z"}},{"metric":{"counter":{"value":44.55},"kind":"absolute","name":"gauge.1","tags":{"host":"my.host.com"},"timestamp":"2021-07-12T07:58:45.223543Z"}}]
...and this configuration...
transforms:
  my_transform_id:
    type: aggregate
    inputs:
    - my-source-or-transform-id
    interval_ms: 5000
[transforms.my_transform_id]
type = "aggregate"
inputs = ["my-source-or-transform-id"]
interval_ms = 5000
{
  "transforms": {
    "my_transform_id": {
      "type": "aggregate",
      "inputs": [
        "my-source-or-transform-id"
      ],
      "interval_ms": 5000
    }
  }
}
...this Vector event is produced:
[{"metric":{"counter":{"value":3.3},"kind":"incremental","name":"counter.1","tags":{"host":"my.host.com"},"timestamp":"2021-07-12T07:58:45.223543Z"}},{"metric":{"counter":{"value":1.1},"kind":"incremental","name":"counter.1","tags":{"host":"different.host.com"},"timestamp":"2021-07-12T07:58:45.223543Z"}},{"metric":{"counter":{"value":44.55},"kind":"absolute","name":"gauge.1","tags":{"host":"my.host.com"},"timestamp":"2021-07-12T07:58:45.223543Z"}}]

How it works

Advantages of Use

The major advantage to aggregation is the reduction of volume. It may reduce costs directly in situations that charge by metric event volume, or indirectly by requiring less CPU to process and/or less network bandwidth to transmit and receive. In systems that are constrained by the processing required to ingest metric events it may help to reduce the processing overhead. This may apply to transforms and sinks downstream of the aggregate transform as well.

Aggregation Behavior

Metrics are aggregated based on their kind. During an interval, incremental metrics are “added” and newer absolute metrics replace older ones in the same series. This results in a reduction of volume and less granularity, while maintaining numerical correctness. As an example, two incremental counter metrics with values 10 and 13 processed by the transform during a period would be aggregated into a single incremental counter with a value of 23. Two absolute gauge metrics with values 93 and 95 would result in a single absolute gauge with the value of 95. More complex types like distribution, histogram, set, and summary behave similarly with incremental values being combined in a manner that makes sense based on their type.

Event-Time Aggregation

When an event_time configuration block is present, metrics are bucketed by the timestamp on each event rather than by the moment Vector processes it. Bucket boundaries are aligned to multiples of interval_ms from the Unix epoch, so the same source timestamp always maps to the same bucket regardless of when Vector receives it. Omit the event_time block to keep the default system-time behavior.

This is useful when downstream sinks key on the metric timestamp. For example, the Datadog Metrics sink overwrites earlier values for an identical timestamp; bucketing on event time prevents distinct samples from collapsing into a single point.

Watermark and Late Events

Vector tracks a watermark — the exclusive end of the most recently emitted bucket. Events whose bucket has already been emitted are dropped and counted via component_discarded_events_total. Use event_time.allowed_lateness_ms to extend how long each bucket accepts events after its window ends (bucket_end + allowed_lateness_ms, compared to the system clock). That cutoff applies when recording an event, not only when the periodic flush runs, so allowed_lateness_ms = 0 enforces strict lateness even if the flush interval is long or misaligned.

Metrics the configured mode does not aggregate (for example an incremental event in mean mode, or an absolute event in sum mode) pass through unchanged, matching system-time behavior, without creating buckets or affecting the watermark. Absolute non-gauge values in mean or stdev mode are ignored (not passed through), also matching system-time behavior.

Missing and Future Timestamps

By default, metrics that will be bucketed and have no timestamp are dropped. Metrics that pass through unchanged (see above) do not require a timestamp. Set event_time.missing_timestamp to use_system_time to fall back to the current system time for bucketed metrics instead. Bucketed events whose timestamp is more than event_time.max_future_ms ahead of the system clock are dropped as a clock-skew guard. All such drops increment component_discarded_events_total (the drop reason is logged, not tagged on the metric).

Shutdown and Reload

When the input stream closes — during shutdown or topology reload — every remaining event-time bucket is flushed before Vector exits, so in-flight metrics are not silently dropped. In diff mode a small rolling window of previous buckets is also retained to compute deltas across bucket boundaries; other modes do not retain previous buckets.

State

This component is stateful, meaning its behavior changes based on previous inputs (events). State is not preserved across restarts, therefore state-dependent behavior will reset between restarts and depend on the inputs (events) received since the most recent restart.