Flow × Function Matrix - Two-Dimensional Classification

Complete navigable/filterable list (stars, dates, licenses, top/flop): site generated from the catalogs → docs/ (GitHub Pages). This document is an editorial analysis.

Objective: Classify solutions along two dimensions: Flow Typology (Streaming/Micro-Batching/Batching) AND Function in the data architecture.


Introduction

This two-dimensional classification allows selecting tools according to:

  1. The temporal axis: Streaming, Micro-Batching, Batching
  2. The functional axis: Collection, Transport, Storage, Processing, Analysis, Governance

The Two Dimensions

Dimension 1: Flow Typology

Mode Latency Main Use Case
Streaming < 1 sec Real-time, alerts, IoT
Micro-Batching 1s - 5min Near real-time dashboards
Batching > 5 min Historical analytics, reports

Dimension 2: Data Function

Function Description Pipeline Position
Collection Capture of source data Start
Transport Routing and delivery Middle
Storage Persistence and organization Base
Processing Transformation and aggregation Middle
Analysis Visualization and BI End
Governance Orchestration, quality, security Cross-cutting

Global Matrix: Flow × Function

Overview

  Collection Transport Storage Processing Analysis Governance
** Streaming** 33 tools 19 tools 20 tools 26 tools 38 tools 9 tools
** Micro-Batching** 24 tools 20 tools 34 tools 45 tools 59 tools 22 tools
** Batching** 12 tools 10 tools 42 tools 35 tools 60 tools 32 tools

Total solutions analyzed: 263


Detailed Classification by Function

FUNCTION: COLLECTION

Initial capture of data from sources (APIs, databases, files, logs, events)

Collection × Flow Matrix

Solution Stream Micro-Batch Batch Category Main strength
Airbyte ○ ◐ ● ETL/ELT Many connectors, incremental sync
Apache Camel ● ● ● Integration 300+ components, versatile
Apache Gobblin ○ ◐ ● Big Data Ingestion Hadoop ecosystem
Apache NiFi ● ● ● Flow-based ETL Visual UI, provenance tracking
Bento ● ● ◐ Stream Processor Declarative config, lightweight
dlt ○ ◐ ● Python Library Schema inference, validation
Meltano ○ ◐ ● ELT Orchestrator Singer taps, DataOps
Singer ○ ◐ ● Standard JSON format, extensible
Starlake (Starflow) ○ ◐ ● Declarative ELT YAML pipelines, multi-warehouse, orchestration (Starflow)
Debezium ● ◐ ○ CDC Real-time binlog capture
Maxwell ● ◐ ○ CDC MySQL CDC to Kafka
Databus ● ◐ ○ CDC LinkedIn low-latency CDC
Fluent Bit ● ● ◐ Log Collector Ultra-lightweight, containers
Fluentd ● ● ◐ Log Collector Unified logging, plugins
Graylog ● ● ● Log Platform Complete platform + SIEM
Logstash ● ● ● Log Pipeline Elastic Stack, filters
Vector ● ● ◐ Observability Rust, high performance
Snowplow ● ● ◐ Event Tracking Analytics pipeline
Rudderstack ● ● ◐ Customer Data Event streaming CDP
Apache SeaTunnel ● ● ● ETL/Integration Multi-engine, massive connectors
CloudQuery ○ ◐ ● Cloud Asset ELT Cloud inventory to SQL
Embulk ○ ◐ ● Bulk Loader Plugins, massive transfers
Apache Flink CDC ● ◐ ○ CDC Real-time capture on Flink
Canal ● ◐ ○ CDC MySQL binlog, Alibaba
PeerDB ● ◐ ○ CDC High-performance Postgres CDC
Sequin ● ◐ ○ CDC Postgres to streams
Grafana Loki ● ● ◐ Log Aggregation Logs indexed by labels
OpenTelemetry Collector ● ● ◐ Telemetry Unified observability standard
Grafana Alloy ● ● ◐ Telemetry Collector OTel/Prometheus collector
rsyslog ● ● ◐ Syslog High-performance log routing
syslog-ng ● ● ◐ Syslog Flexible log collection

Collection Leaders: (excluding new additions)


FUNCTION: TRANSPORT

Reliable delivery and routing of data between systems

Transport × Flow Matrix

Solution Stream Micro-Batch Batch Category Main strength
Apache Kafka ● ● ◐ Event Streaming Market leader, ecosystem
Apache Pulsar ● ● ◐ Messaging Platform Multi-tenancy, geo-replication
NATS ● ◐ ○ Messaging Ultra-lightweight, microsecond latency
RabbitMQ ● ● ◐ Message Broker AMQP, messaging patterns
Redpanda ● ● ◐ Kafka-compatible C++, high performance
Apache Camel ● ● ● Integration Routes, transformations
Apache NiFi ● ● ● Data Flow Back-pressure, priorities
Bento ● ● ◐ Stream Processor Declarative pipelines
Apache RocketMQ ● ● ◐ Message Broker Transactions, Alibaba scalability
Eclipse Mosquitto ● ◐ ○ MQTT Broker Lightweight, IoT edge
EMQX ● ● ◐ MQTT Broker Massively distributed, IoT
Apache Pekko ● ● ◐ Actor/Messaging Reactive actors, Akka fork

Transport Leaders: (excluding new additions)


FUNCTION: STORAGE

Persistence, organization and access to data

Storage × Flow Matrix

Solution Stream Micro-Batch Batch Category Main strength
Avro ● ● ● Serialization Schema evolution, compact
ORC ○ ◐ ● Columnar Format Hive-optimized, compression
Parquet ○ ◐ ● Columnar Format Standard analytics, Spark
MinIO ◐ ● ● Object Storage S3-compatible, data lake
Delta Lake ◐ ● ● Table Format ACID, time travel, Spark
Hudi ● ● ● Table Format Streaming upserts, versatile
Iceberg ◐ ● ● Table Format Schema evolution, partitioning
Paimon ● ● ● Table Format Streaming lake, Flink
Nessie ○ ◐ ● Catalog Git-like versioning
Polaris ○ ◐ ● Catalog Unified metadata
Hive Metastore ○ ◐ ● Metastore Hadoop metadata
ClickHouse ● ● ● OLAP Database Real-time ingestion, analytics
PostgreSQL ● ● ● RDBMS OLTP, streaming extensions
Apache Arrow ● ● ● In-Memory Format Zero-copy columnar, interop
Apache Kudu ● ● ● Columnar Store Updates + analytical scans
Lance ○ ◐ ● Columnar Format ML/vector format, random access
DuckLake ○ ◐ ● Lakehouse Format SQL catalog for lakehouse
Vortex ○ ◐ ● Columnar Format Next-generation compression
Ceph ◐ ● ● Object Storage Unified distributed storage
SeaweedFS ◐ ● ● Object Storage Highly scalable blob/files
JuiceFS ◐ ● ● Distributed FS POSIX over object storage
Garage ◐ ● ● Object Storage Lightweight S3, geo-distributed
Apache Ozone ◐ ● ● Object Storage Scalable, Hadoop-compatible
lakeFS ○ ◐ ● Data Versioning Git-like for data lakes
Unity Catalog ○ ◐ ● Catalog Open catalog, governance
Apache Gravitino ○ ◐ ● Metadata Catalog Multi-source metadata
Lakekeeper ○ ◐ ● Iceberg Catalog Iceberg REST catalog in Rust
Apache XTable ○ ◐ ● Format Interop Delta/Hudi/Iceberg conversion
Milvus ● ● ● Vector DB Scalable vector search
Qdrant ● ● ● Vector DB Rust, rich filters
Weaviate ● ● ● Vector DB Hybrid, ML modules
Chroma ● ● ● Vector DB Embedded, dev-friendly
pgvector ● ● ● Vector Extension Vectors inside Postgres
GreptimeDB ● ● ● Time-Series DB Unified metrics/logs/traces
QuestDB ● ● ● Time-Series DB Fast ingestion, SQL
TimescaleDB ● ● ● Time-Series DB Postgres extension, hypertables
InfluxDB ● ● ● Time-Series DB Metrics standard, IoT
Apache IoTDB ● ● ● Time-Series DB Industrial IoT, edge/cloud
VictoriaMetrics ● ● ● Time-Series DB Scalable metrics, Prometheus
TDengine ● ● ● Time-Series DB High-frequency IoT
CrateDB ● ● ● Distributed SQL Real-time distributed SQL

Storage Leaders: (excluding new additions)


FUNCTION: PROCESSING

Transformation, aggregation and enrichment of data

Processing × Flow Matrix

Solution Stream Micro-Batch Batch Category Main strength
Apache Flink ● ● ◐ Stream Processor True streaming, stateful
Apache Storm ● ◐ ○ Stream Processor Real-time computation
Apache Samza ● ● ◐ Stream Processor Kafka-based, stateful
Materialize ● ● ◐ Streaming SQL Incremental views
Apache Spark ◐ ● ● Unified Engine Structured Streaming, batch
Apache Beam ● ● ● Unified API Portability, multi-runner
Apache Hop ○ ◐ ● ETL Platform Visual workflows
dbt core ○ ◐ ● SQL Transform Analytics engineering
Pandas ○ ◐ ● Dataframe In-memory, Python
Polars ○ ● ● Dataframe Rust, lazy evaluation
Dask ○ ● ● Parallel Computing Scale Pandas
Ibis ○ ◐ ● Dataframe API Portable, lazy
DuckDB ○ ● ● OLAP Engine Embedded, fast analytics
Quack on Demand ○ ● ● DuckDB gateway Autoscaling DuckDB fleet, federated queries
Trino ○ ◐ ● Query Engine MPP SQL, data lake
Presto ○ ◐ ● Query Engine Interactive queries
Druid ● ● ◐ OLAP Real-time aggregations
Pinot ● ● ◐ OLAP Sub-second queries
StarRocks ◐ ● ● MPP Database Unified analytics
Doris ◐ ● ● MPP Database Real-time + batch
chDB ○ ● ● Embedded OLAP ClickHouse in-process
Velox ● ● ● Execution Engine Reusable vectorized engine
RisingWave ● ● ◐ Streaming SQL Streaming materialized views
Arroyo ● ● ○ Stream Processor SQL streaming in Rust
Timeplus Proton ● ● ◐ Streaming SQL Unified streaming analytics
Bytewax ● ● ○ Stream Processor Native Python streaming
Quix Streams ● ● ○ Stream Processor Kafka streaming in Python
Faust-streaming ● ● ○ Stream Processor Python streaming, Kafka
Pathway ● ● ◐ Stream Processor Real-time ETL + LLM
Numaflow ● ● ○ Stream Processor Kubernetes-native streaming
Apache StreamPipes ● ● ◐ IIoT Processing Industrial stream toolbox
Feldera ● ● ◐ Incremental Compute Continuous incremental compute
Apache Sedona ○ ● ● Geospatial Geospatial big data on Spark
Duckle ○ ● ● DuckDB Tooling Tooling around DuckDB
Fugue ○ ● ● Abstraction Layer Unified code Spark/Dask/Ray
Ray ◐ ● ● Distributed Compute Distributed compute Python/ML
Odyssée ○ ◐ ● Data Processing Data processing
Daft ○ ● ● Dataframe Distributed multimodal dataframe
Modin ○ ● ● Dataframe Parallelized Pandas
cuDF (RAPIDS) ○ ● ● GPU Dataframe GPU-accelerated Pandas
Xarray ○ ◐ ● N-D Arrays Labeled scientific arrays
Vaex ○ ● ● Dataframe Out-of-core, lazy

Processing Leaders: (excluding new additions)


FUNCTION: ANALYSIS

Visualization, exploration and business intelligence

Analysis × Flow Matrix

Solution Stream Micro-Batch Batch Category Main strength
D3JS ● ● ● Framework Maximum flexibility, web
Plotly ● ● ● Framework Interactive, Dash
Apache ECharts ● ● ● Framework Enterprise viz, performant
Chart JS ● ● ● Framework Simple, lightweight
Bokeh ● ● ● Framework Python viz, callbacks
Matplotlib ○ ◐ ● Framework Static viz, Python
Seaborn ○ ◐ ● Framework Statistical viz
Streamlit ● ● ● High-Code Rapid prototyping, Python
Plotly Dash ● ● ● High-Code Production dashboards
Panel ● ● ● High-Code HoloViz, versatile
Taipy ● ● ● High-Code Pipelines + GUI
Grafana ● ● ● Low-Code Monitoring, time-series
Kibana ● ● ● Low-Code Elastic Stack, logs
PyGWalker ○ ◐ ● Low-Code Tableau-like, Jupyter
Superset ○ ● ● No-Code BI platform, SQL
Metabase ○ ● ● No-Code Easy BI, auto-refresh
Lightdash ○ ● ● No-Code dbt-native BI
Matomo ● ● ● Web Analytics Privacy-first, real-time
Plausible ● ● ◐ Web Analytics Lightweight analytics
Posthog ● ● ● Web Analytics Product analytics, events
AntV G2 ● ● ● Framework Graphical grammar, web
deck.gl ● ● ● Geo Viz Large-scale WebGL geospatial
Great Tables ○ ◐ ● Table Viz Polished tables in Python
kepler.gl ◐ ● ● Geo Viz Exploratory geospatial
Lonboard ◐ ● ● Geo Viz Vector maps in Python
Nivo ● ● ● Framework React viz components
Perspective ● ● ● Streaming Viz Real-time tables/charts
Recharts ● ● ● Framework Declarative React charts
visx ● ● ● Framework React/D3 viz primitives
ApexCharts ● ● ● Framework Interactive JS charts
ggplot2 ○ ◐ ● Framework Graphical grammar in R
Plotnine ○ ◐ ● Framework ggplot for Python
Gradio ● ● ● High-Code ML UI/quick demos
Marimo ● ● ● High-Code Reactive Python notebooks
NiceGUI ● ● ● High-Code Simple Python web UI
Quarto ○ ◐ ● Publishing Technical documents/reports
Reflex ● ● ● High-Code Full-Python web apps
Solara ● ● ● High-Code Reactive Python apps
Voila ● ● ● High-Code Notebooks as web apps
H2O Wave ● ● ● High-Code Real-time AI/ML apps
Shiny (R) ● ● ● High-Code Interactive web apps in R
Datasette ○ ◐ ● Data Publishing SQLite exploration/publishing
Observable Framework ● ● ● Publishing Static data dashboards
Vizro ○ ● ● Low-Code Configurable Plotly dashboards
Querybook ○ ● ● SQL Notebook Collaborative SQL notebooks
WrenAI ○ ● ● BI/Text-to-SQL Natural language queries
Appsmith ◐ ● ● Low-Code App Internal data tools
Budibase ◐ ● ● Low-Code App Fast internal apps
DataEase ○ ● ● No-Code BI Open source BI dashboards
ToolJet ◐ ● ● Low-Code App Internal apps/dashboards
Ackee ● ● ◐ Web Analytics Privacy-respecting analytics
GoatCounter ● ● ◐ Web Analytics Lightweight, cookieless
Open Web Analytics ● ● ● Web Analytics Self-hosted alternative
OpenReplay ● ● ◐ Session Replay Replay and UX debugging
Umami ● ● ◐ Web Analytics Simple, privacy-first
Rybbit ● ● ◐ Web Analytics Modern open source analytics

Analysis Leaders: (excluding new additions)


FUNCTION: GOVERNANCE

Orchestration, quality, security, compliance

Governance × Flow Matrix

Solution Stream Micro-Batch Batch Category Main strength
Apache Airflow ○ ◐ ● Orchestration Batch leader, DAG
Dagster ○ ◐ ● Orchestration Asset-based, data-aware
Prefect ○ ◐ ● Orchestration Dynamic workflows
Kestra ◐ ● ● Orchestration Event-driven, versatile
DataHub ○ ◐ ● Data Catalog Metadata discovery
OpenMetadata ○ ◐ ● Data Catalog Open standard, lineage
Great Expectations ○ ● ● Data Quality Validation, profiling
SQLFluff ○ ○ ● SQL Linter Code quality
Pandera ○ ● ● Data Quality Python dataframe validation
Evidently ○ ● ● ML Monitoring ML drift and quality
OpenLineage ◐ ● ● Data Lineage Open lineage standard
CKAN ○ ◐ ● Open Data Portal Open data catalog
Apache Ranger ◐ ● ● Security Centralized access policies
Apache Egeria ○ ◐ ● Metadata Governance Cross-enterprise metadata
Node-RED ● ● ◐ Flow Automation Visual IoT automation
Activepieces ◐ ● ● Automation No-code workflow automation
Huginn ◐ ● ● Automation Self-hosted agents and alerts
Kepler ○ ● ● Energy Monitoring Container energy consumption
Scaphandre ● ● ◐ Energy Monitoring Real-time energy metrics
Flyte ○ ◐ ● Orchestration Scalable ML/data workflows
Temporal ◐ ● ● Workflow Engine Reliable durable workflows
Kedro ○ ◐ ● Pipeline Framework Reproducible data pipelines
Metaflow ○ ◐ ● ML Pipeline Netflix ML workflows
ZenML ○ ◐ ● MLOps Portable ML pipelines
Windmill ◐ ● ● Workflow/Automation Dev scripts and workflows
Hamilton ○ ◐ ● Pipeline Framework Dataflow function DAG
Astronomer Cosmos ○ ◐ ● Orchestration dbt inside Airflow
Microsoft Presidio ○ ● ● PII/Privacy PII detection/anonymization
Faker ○ ◐ ● Synthetic Data Fake data generation
OPA ● ● ● Policy Engine Policies as code (Rego)
OpenFGA ● ● ● Authorization Fine-grained authorization (ReBAC)
Cerbos ● ● ● Authorization Decoupled authorization
SDV ○ ◐ ● Synthetic Data Realistic synthetic data
Permify ● ● ● Authorization Google Zanzibar authorization

Governance Leaders: (excluding new additions)


Recommendation Matrices

Matrix 1: By Use Case × Function

Use Case Collection Transport Storage Processing Analysis Governance
** Real-Time Alerts** Debezium Kafka Hudi Flink Grafana Kestra
** Real-Time Dashboard** Vector Kafka ClickHouse Flink Kibana -
** Fraud Detection** CDC Kafka Druid Flink Dash -
** Near-Real-Time BI** NiFi Pulsar Delta Lake Spark Superset Airflow
** Data Warehouse** Airbyte - Iceberg dbt Superset Airflow
** ML Pipeline** dlt - Parquet Spark Jupyter Airflow
** IoT Platform** MQTT NATS Paimon Flink Grafana Kestra
** Web Analytics** Snowplow Kafka ClickHouse Druid Posthog -

Matrix 2: Full Stack by Flow × All

PURE STREAMING STACK (Latency < 1s)

Function Recommended Solution Alternative
Collection Debezium (CDC) Maxwell, Fluent Bit
Transport Apache Kafka NATS, Pulsar
Storage Hudi Paimon, ClickHouse
Processing Apache Flink Storm, Materialize
Analysis Grafana Kibana, Dash
Governance Kestra -

Architecture Example:

MySQL → Debezium → Kafka → Flink → Hudi → Trino → Grafana
                             ↓
                         ClickHouse → Superset

MICRO-BATCHING STACK (Latency 1s-5min)

Function Recommended Solution Alternative
Collection Apache NiFi Camel, Bento
Transport Apache Kafka Pulsar
Storage Delta Lake Iceberg
Processing Apache Spark Beam, DuckDB
Analysis Superset Metabase
Governance Airflow Dagster

Architecture Example:

APIs → NiFi → Kafka → Spark Streaming → Delta Lake → Trino → Superset
                                            ↓
                                        StarRocks → Grafana (refresh 1min)

BATCHING STACK (Latency > 5min)

Function Recommended Solution Alternative
Collection Airbyte Meltano, dlt
Transport - NiFi (if needed)
Storage Iceberg Parquet, Delta Lake
Processing dbt + Spark Trino, Hop
Analysis Superset Metabase, Lightdash
Governance Airflow Dagster, Prefect

Architecture Example:

Sources → Airbyte → MinIO/S3 (Raw) → Spark → Iceberg (Processed)
                                       ↓
                                      dbt → Iceberg (Curated) → Trino → Superset
                                                                           ↓
                                                                      Airflow (orchestration)

Summary Tables

Top 10 Multi-Function Solutions

Solution Functions Covered Flow Modes Versatility Score
Apache NiFi Collection + Transport + Governance ●●●  
Apache Kafka Transport + Processing ●●◐  
Apache Spark Processing + Storage ◐●●  
ClickHouse Storage + Processing ●●●  
Grafana Analysis + Monitoring ●●●  
Apache Flink Streaming processing ●●◐  
Hudi Versatile storage ●●●  
Superset BI analysis ○●●  
Airflow Batch orchestration ○◐●  
dbt core SQL transform ○◐●  

Functional Coverage by Flow

Flow Collection Transport Storage Processing Analysis Governance Total
** Streaming** 33 19 20 26 38 9 145
** Micro-Batch** 24 20 34 45 59 22 204
** Batching** 12 10 42 35 60 32 191

Insight: Micro-batching offers the best overall coverage (204 tool-function combinations).


Gap Analysis by Function

Function Streaming Coverage Batch Coverage Main Gap
Collection [OK] Excellent [OK] Excellent -
Transport [OK] Excellent [!] Limited Batch does not require transport
Storage [!] Limited [OK] Excellent Limited streaming formats
Processing [OK] Good [OK] Excellent -
Analysis [OK] Good [OK] Excellent -
Governance [X] Weak [OK] Excellent Missing streaming orchestration

Recommendations:


Two-Dimensional Selection Guide

Step 1: Identify the Function

Question: What is the position in the pipeline?

Step 2: Identify the Flow

Question: What is the maximum acceptable latency?

Step 3: Consult the Matrix

Refer to the corresponding section (e.g., “Collection × Streaming”)

Decision Example

Need: E-commerce dashboards with real-time metrics (sales, inventory)

  1. Function: Analysis (end of pipeline)
  2. Required latency: 10 seconds → Micro-Batching
  3. Analysis × Micro-Batch Matrix:
    • Options: Superset ●, Grafana ●, Metabase ●
  4. Choice: Superset (rich BI) + Grafana (monitoring)

Two-Dimensional Architecture Patterns

Pattern 1: Full Lambda (All Functions)

 COLLECTION
├─ Streaming: Debezium (CDC)
└─ Batching: Airbyte (Bulk)

 TRANSPORT
└─ Streaming: Kafka

 STORAGE
├─ Hot: ClickHouse (streaming)
└─ Cold: Iceberg (batch)

 PROCESSING
├─ Streaming: Flink
└─ Batching: Spark + dbt

 ANALYSIS
├─ Real-time: Grafana
└─ BI: Superset

 GOVERNANCE
└─ Batching: Airflow

Pattern 2: Simplified Kappa (Streaming Only)

 COLLECTION: Fluent Bit
 TRANSPORT: NATS
 STORAGE: Paimon
 PROCESSING: Flink
 ANALYSIS: Grafana
 GOVERNANCE: Kestra (event-driven)

Pattern 3: Modern Data Stack (Batch + BI)

 COLLECTION: Airbyte
 STORAGE: Iceberg on MinIO
 PROCESSING: dbt
 ANALYSIS: Superset
 GOVERNANCE: Airflow + DataHub

Heatmap: Solution Maturity

Maturity Legend

  Collection Transport Storage Processing Analysis Governance
** Stream** ● ● ◐ ● ● ○
** Micro** ● ● ● ● ● ◐
** Batch** ● ● ● ● ● ●

Insights:


Cross-References


Classification Methodology

Function Criteria

Each tool is classified according to its primary function:

  1. Collection: Initial capture from sources
  2. Transport: Reliable delivery between systems
  3. Storage: Persistence and organization
  4. Processing: Transformation and enrichment
  5. Analysis: Visualization and BI
  6. Governance: Orchestration and quality

Flow Criteria

Classification according to 3 modes based on latency observed in production:

Validation

Each classification is based on:


Document created on: 2025-12-09 Last updated: 2026-06-25 Version: 2.1 Status: [OK] Complete - 263 solutions × 6 functions × 3 flows