Flow × Function Matrix - Two-Dimensional Classification
Complete navigable/filterable list (stars, dates, licenses, top/flop): site generated from the catalogs →
docs/(GitHub Pages). This document is an editorial analysis.
Objective: Classify solutions along two dimensions: Flow Typology (Streaming/Micro-Batching/Batching) AND Function in the data architecture.
Introduction
This two-dimensional classification allows selecting tools according to:
- The temporal axis: Streaming, Micro-Batching, Batching
- The functional axis: Collection, Transport, Storage, Processing, Analysis, Governance
The Two Dimensions
Dimension 1: Flow Typology
| Mode | Latency | Main Use Case |
|---|---|---|
| Streaming | < 1 sec | Real-time, alerts, IoT |
| Micro-Batching | 1s - 5min | Near real-time dashboards |
| Batching | > 5 min | Historical analytics, reports |
Dimension 2: Data Function
| Function | Description | Pipeline Position |
|---|---|---|
| Collection | Capture of source data | Start |
| Transport | Routing and delivery | Middle |
| Storage | Persistence and organization | Base |
| Processing | Transformation and aggregation | Middle |
| Analysis | Visualization and BI | End |
| Governance | Orchestration, quality, security | Cross-cutting |
Global Matrix: Flow × Function
Overview
| Collection | Transport | Storage | Processing | Analysis | Governance | |
|---|---|---|---|---|---|---|
| ** Streaming** | 33 tools | 19 tools | 20 tools | 26 tools | 38 tools | 9 tools |
| ** Micro-Batching** | 24 tools | 20 tools | 34 tools | 45 tools | 59 tools | 22 tools |
| ** Batching** | 12 tools | 10 tools | 42 tools | 35 tools | 60 tools | 32 tools |
Total solutions analyzed: 263
Detailed Classification by Function
FUNCTION: COLLECTION
Initial capture of data from sources (APIs, databases, files, logs, events)
Collection × Flow Matrix
| Solution | Stream | Micro-Batch | Batch | Category | Main strength |
|---|---|---|---|---|---|
| Airbyte | ○ | ◐ | ● | ETL/ELT | Many connectors, incremental sync |
| Apache Camel | ● | ● | ● | Integration | 300+ components, versatile |
| Apache Gobblin | ○ | ◐ | ● | Big Data Ingestion | Hadoop ecosystem |
| Apache NiFi | ● | ● | ● | Flow-based ETL | Visual UI, provenance tracking |
| Bento | ● | ● | ◐ | Stream Processor | Declarative config, lightweight |
| dlt | ○ | ◐ | ● | Python Library | Schema inference, validation |
| Meltano | ○ | ◐ | ● | ELT Orchestrator | Singer taps, DataOps |
| Singer | ○ | ◐ | ● | Standard | JSON format, extensible |
| Starlake (Starflow) | ○ | ◐ | ● | Declarative ELT | YAML pipelines, multi-warehouse, orchestration (Starflow) |
| Debezium | ● | ◐ | ○ | CDC | Real-time binlog capture |
| Maxwell | ● | ◐ | ○ | CDC | MySQL CDC to Kafka |
| Databus | ● | ◐ | ○ | CDC | LinkedIn low-latency CDC |
| Fluent Bit | ● | ● | ◐ | Log Collector | Ultra-lightweight, containers |
| Fluentd | ● | ● | ◐ | Log Collector | Unified logging, plugins |
| Graylog | ● | ● | ● | Log Platform | Complete platform + SIEM |
| Logstash | ● | ● | ● | Log Pipeline | Elastic Stack, filters |
| Vector | ● | ● | ◐ | Observability | Rust, high performance |
| Snowplow | ● | ● | ◐ | Event Tracking | Analytics pipeline |
| Rudderstack | ● | ● | ◐ | Customer Data | Event streaming CDP |
| Apache SeaTunnel | ● | ● | ● | ETL/Integration | Multi-engine, massive connectors |
| CloudQuery | ○ | ◐ | ● | Cloud Asset ELT | Cloud inventory to SQL |
| Embulk | ○ | ◐ | ● | Bulk Loader | Plugins, massive transfers |
| Apache Flink CDC | ● | ◐ | ○ | CDC | Real-time capture on Flink |
| Canal | ● | ◐ | ○ | CDC | MySQL binlog, Alibaba |
| PeerDB | ● | ◐ | ○ | CDC | High-performance Postgres CDC |
| Sequin | ● | ◐ | ○ | CDC | Postgres to streams |
| Grafana Loki | ● | ● | ◐ | Log Aggregation | Logs indexed by labels |
| OpenTelemetry Collector | ● | ● | ◐ | Telemetry | Unified observability standard |
| Grafana Alloy | ● | ● | ◐ | Telemetry Collector | OTel/Prometheus collector |
| rsyslog | ● | ● | ◐ | Syslog | High-performance log routing |
| syslog-ng | ● | ● | ◐ | Syslog | Flexible log collection |
Collection Leaders: (excluding new additions)
- Streaming: Debezium (CDC), Vector (logs), NiFi (versatile)
- Micro-Batching: NiFi, Camel, Graylog
- Batching: Airbyte (ETL), Meltano, dlt
FUNCTION: TRANSPORT
Reliable delivery and routing of data between systems
Transport × Flow Matrix
| Solution | Stream | Micro-Batch | Batch | Category | Main strength |
|---|---|---|---|---|---|
| Apache Kafka | ● | ● | ◐ | Event Streaming | Market leader, ecosystem |
| Apache Pulsar | ● | ● | ◐ | Messaging Platform | Multi-tenancy, geo-replication |
| NATS | ● | ◐ | ○ | Messaging | Ultra-lightweight, microsecond latency |
| RabbitMQ | ● | ● | ◐ | Message Broker | AMQP, messaging patterns |
| Redpanda | ● | ● | ◐ | Kafka-compatible | C++, high performance |
| Apache Camel | ● | ● | ● | Integration | Routes, transformations |
| Apache NiFi | ● | ● | ● | Data Flow | Back-pressure, priorities |
| Bento | ● | ● | ◐ | Stream Processor | Declarative pipelines |
| Apache RocketMQ | ● | ● | ◐ | Message Broker | Transactions, Alibaba scalability |
| Eclipse Mosquitto | ● | ◐ | ○ | MQTT Broker | Lightweight, IoT edge |
| EMQX | ● | ● | ◐ | MQTT Broker | Massively distributed, IoT |
| Apache Pekko | ● | ● | ◐ | Actor/Messaging | Reactive actors, Akka fork |
Transport Leaders: (excluding new additions)
- Streaming: Kafka (standard), NATS (lightweight), Pulsar (multi-DC)
- Micro-Batching: Kafka, Pulsar, RabbitMQ
- Batching: NiFi, Camel
FUNCTION: STORAGE
Persistence, organization and access to data
Storage × Flow Matrix
| Solution | Stream | Micro-Batch | Batch | Category | Main strength |
|---|---|---|---|---|---|
| Avro | ● | ● | ● | Serialization | Schema evolution, compact |
| ORC | ○ | ◐ | ● | Columnar Format | Hive-optimized, compression |
| Parquet | ○ | ◐ | ● | Columnar Format | Standard analytics, Spark |
| MinIO | ◐ | ● | ● | Object Storage | S3-compatible, data lake |
| Delta Lake | ◐ | ● | ● | Table Format | ACID, time travel, Spark |
| Hudi | ● | ● | ● | Table Format | Streaming upserts, versatile |
| Iceberg | ◐ | ● | ● | Table Format | Schema evolution, partitioning |
| Paimon | ● | ● | ● | Table Format | Streaming lake, Flink |
| Nessie | ○ | ◐ | ● | Catalog | Git-like versioning |
| Polaris | ○ | ◐ | ● | Catalog | Unified metadata |
| Hive Metastore | ○ | ◐ | ● | Metastore | Hadoop metadata |
| ClickHouse | ● | ● | ● | OLAP Database | Real-time ingestion, analytics |
| PostgreSQL | ● | ● | ● | RDBMS | OLTP, streaming extensions |
| Apache Arrow | ● | ● | ● | In-Memory Format | Zero-copy columnar, interop |
| Apache Kudu | ● | ● | ● | Columnar Store | Updates + analytical scans |
| Lance | ○ | ◐ | ● | Columnar Format | ML/vector format, random access |
| DuckLake | ○ | ◐ | ● | Lakehouse Format | SQL catalog for lakehouse |
| Vortex | ○ | ◐ | ● | Columnar Format | Next-generation compression |
| Ceph | ◐ | ● | ● | Object Storage | Unified distributed storage |
| SeaweedFS | ◐ | ● | ● | Object Storage | Highly scalable blob/files |
| JuiceFS | ◐ | ● | ● | Distributed FS | POSIX over object storage |
| Garage | ◐ | ● | ● | Object Storage | Lightweight S3, geo-distributed |
| Apache Ozone | ◐ | ● | ● | Object Storage | Scalable, Hadoop-compatible |
| lakeFS | ○ | ◐ | ● | Data Versioning | Git-like for data lakes |
| Unity Catalog | ○ | ◐ | ● | Catalog | Open catalog, governance |
| Apache Gravitino | ○ | ◐ | ● | Metadata Catalog | Multi-source metadata |
| Lakekeeper | ○ | ◐ | ● | Iceberg Catalog | Iceberg REST catalog in Rust |
| Apache XTable | ○ | ◐ | ● | Format Interop | Delta/Hudi/Iceberg conversion |
| Milvus | ● | ● | ● | Vector DB | Scalable vector search |
| Qdrant | ● | ● | ● | Vector DB | Rust, rich filters |
| Weaviate | ● | ● | ● | Vector DB | Hybrid, ML modules |
| Chroma | ● | ● | ● | Vector DB | Embedded, dev-friendly |
| pgvector | ● | ● | ● | Vector Extension | Vectors inside Postgres |
| GreptimeDB | ● | ● | ● | Time-Series DB | Unified metrics/logs/traces |
| QuestDB | ● | ● | ● | Time-Series DB | Fast ingestion, SQL |
| TimescaleDB | ● | ● | ● | Time-Series DB | Postgres extension, hypertables |
| InfluxDB | ● | ● | ● | Time-Series DB | Metrics standard, IoT |
| Apache IoTDB | ● | ● | ● | Time-Series DB | Industrial IoT, edge/cloud |
| VictoriaMetrics | ● | ● | ● | Time-Series DB | Scalable metrics, Prometheus |
| TDengine | ● | ● | ● | Time-Series DB | High-frequency IoT |
| CrateDB | ● | ● | ● | Distributed SQL | Real-time distributed SQL |
Storage Leaders: (excluding new additions)
- Streaming: Hudi, Paimon, ClickHouse
- Micro-Batching: Delta Lake, Iceberg, Hudi, MinIO
- Batching: Parquet, ORC, Iceberg, MinIO
FUNCTION: PROCESSING
Transformation, aggregation and enrichment of data
Processing × Flow Matrix
| Solution | Stream | Micro-Batch | Batch | Category | Main strength |
|---|---|---|---|---|---|
| Apache Flink | ● | ● | ◐ | Stream Processor | True streaming, stateful |
| Apache Storm | ● | ◐ | ○ | Stream Processor | Real-time computation |
| Apache Samza | ● | ● | ◐ | Stream Processor | Kafka-based, stateful |
| Materialize | ● | ● | ◐ | Streaming SQL | Incremental views |
| Apache Spark | ◐ | ● | ● | Unified Engine | Structured Streaming, batch |
| Apache Beam | ● | ● | ● | Unified API | Portability, multi-runner |
| Apache Hop | ○ | ◐ | ● | ETL Platform | Visual workflows |
| dbt core | ○ | ◐ | ● | SQL Transform | Analytics engineering |
| Pandas | ○ | ◐ | ● | Dataframe | In-memory, Python |
| Polars | ○ | ● | ● | Dataframe | Rust, lazy evaluation |
| Dask | ○ | ● | ● | Parallel Computing | Scale Pandas |
| Ibis | ○ | ◐ | ● | Dataframe API | Portable, lazy |
| DuckDB | ○ | ● | ● | OLAP Engine | Embedded, fast analytics |
| Quack on Demand | ○ | ● | ● | DuckDB gateway | Autoscaling DuckDB fleet, federated queries |
| Trino | ○ | ◐ | ● | Query Engine | MPP SQL, data lake |
| Presto | ○ | ◐ | ● | Query Engine | Interactive queries |
| Druid | ● | ● | ◐ | OLAP | Real-time aggregations |
| Pinot | ● | ● | ◐ | OLAP | Sub-second queries |
| StarRocks | ◐ | ● | ● | MPP Database | Unified analytics |
| Doris | ◐ | ● | ● | MPP Database | Real-time + batch |
| chDB | ○ | ● | ● | Embedded OLAP | ClickHouse in-process |
| Velox | ● | ● | ● | Execution Engine | Reusable vectorized engine |
| RisingWave | ● | ● | ◐ | Streaming SQL | Streaming materialized views |
| Arroyo | ● | ● | ○ | Stream Processor | SQL streaming in Rust |
| Timeplus Proton | ● | ● | ◐ | Streaming SQL | Unified streaming analytics |
| Bytewax | ● | ● | ○ | Stream Processor | Native Python streaming |
| Quix Streams | ● | ● | ○ | Stream Processor | Kafka streaming in Python |
| Faust-streaming | ● | ● | ○ | Stream Processor | Python streaming, Kafka |
| Pathway | ● | ● | ◐ | Stream Processor | Real-time ETL + LLM |
| Numaflow | ● | ● | ○ | Stream Processor | Kubernetes-native streaming |
| Apache StreamPipes | ● | ● | ◐ | IIoT Processing | Industrial stream toolbox |
| Feldera | ● | ● | ◐ | Incremental Compute | Continuous incremental compute |
| Apache Sedona | ○ | ● | ● | Geospatial | Geospatial big data on Spark |
| Duckle | ○ | ● | ● | DuckDB Tooling | Tooling around DuckDB |
| Fugue | ○ | ● | ● | Abstraction Layer | Unified code Spark/Dask/Ray |
| Ray | ◐ | ● | ● | Distributed Compute | Distributed compute Python/ML |
| Odyssée | ○ | ◐ | ● | Data Processing | Data processing |
| Daft | ○ | ● | ● | Dataframe | Distributed multimodal dataframe |
| Modin | ○ | ● | ● | Dataframe | Parallelized Pandas |
| cuDF (RAPIDS) | ○ | ● | ● | GPU Dataframe | GPU-accelerated Pandas |
| Xarray | ○ | ◐ | ● | N-D Arrays | Labeled scientific arrays |
| Vaex | ○ | ● | ● | Dataframe | Out-of-core, lazy |
Processing Leaders: (excluding new additions)
- Streaming: Flink (leader), Materialize, Druid, Pinot
- Micro-Batching: Spark, Beam, DuckDB, Polars
- Batching: Spark, Trino, dbt, Presto
FUNCTION: ANALYSIS
Visualization, exploration and business intelligence
Analysis × Flow Matrix
| Solution | Stream | Micro-Batch | Batch | Category | Main strength |
|---|---|---|---|---|---|
| D3JS | ● | ● | ● | Framework | Maximum flexibility, web |
| Plotly | ● | ● | ● | Framework | Interactive, Dash |
| Apache ECharts | ● | ● | ● | Framework | Enterprise viz, performant |
| Chart JS | ● | ● | ● | Framework | Simple, lightweight |
| Bokeh | ● | ● | ● | Framework | Python viz, callbacks |
| Matplotlib | ○ | ◐ | ● | Framework | Static viz, Python |
| Seaborn | ○ | ◐ | ● | Framework | Statistical viz |
| Streamlit | ● | ● | ● | High-Code | Rapid prototyping, Python |
| Plotly Dash | ● | ● | ● | High-Code | Production dashboards |
| Panel | ● | ● | ● | High-Code | HoloViz, versatile |
| Taipy | ● | ● | ● | High-Code | Pipelines + GUI |
| Grafana | ● | ● | ● | Low-Code | Monitoring, time-series |
| Kibana | ● | ● | ● | Low-Code | Elastic Stack, logs |
| PyGWalker | ○ | ◐ | ● | Low-Code | Tableau-like, Jupyter |
| Superset | ○ | ● | ● | No-Code | BI platform, SQL |
| Metabase | ○ | ● | ● | No-Code | Easy BI, auto-refresh |
| Lightdash | ○ | ● | ● | No-Code | dbt-native BI |
| Matomo | ● | ● | ● | Web Analytics | Privacy-first, real-time |
| Plausible | ● | ● | ◐ | Web Analytics | Lightweight analytics |
| Posthog | ● | ● | ● | Web Analytics | Product analytics, events |
| AntV G2 | ● | ● | ● | Framework | Graphical grammar, web |
| deck.gl | ● | ● | ● | Geo Viz | Large-scale WebGL geospatial |
| Great Tables | ○ | ◐ | ● | Table Viz | Polished tables in Python |
| kepler.gl | ◐ | ● | ● | Geo Viz | Exploratory geospatial |
| Lonboard | ◐ | ● | ● | Geo Viz | Vector maps in Python |
| Nivo | ● | ● | ● | Framework | React viz components |
| Perspective | ● | ● | ● | Streaming Viz | Real-time tables/charts |
| Recharts | ● | ● | ● | Framework | Declarative React charts |
| visx | ● | ● | ● | Framework | React/D3 viz primitives |
| ApexCharts | ● | ● | ● | Framework | Interactive JS charts |
| ggplot2 | ○ | ◐ | ● | Framework | Graphical grammar in R |
| Plotnine | ○ | ◐ | ● | Framework | ggplot for Python |
| Gradio | ● | ● | ● | High-Code | ML UI/quick demos |
| Marimo | ● | ● | ● | High-Code | Reactive Python notebooks |
| NiceGUI | ● | ● | ● | High-Code | Simple Python web UI |
| Quarto | ○ | ◐ | ● | Publishing | Technical documents/reports |
| Reflex | ● | ● | ● | High-Code | Full-Python web apps |
| Solara | ● | ● | ● | High-Code | Reactive Python apps |
| Voila | ● | ● | ● | High-Code | Notebooks as web apps |
| H2O Wave | ● | ● | ● | High-Code | Real-time AI/ML apps |
| Shiny (R) | ● | ● | ● | High-Code | Interactive web apps in R |
| Datasette | ○ | ◐ | ● | Data Publishing | SQLite exploration/publishing |
| Observable Framework | ● | ● | ● | Publishing | Static data dashboards |
| Vizro | ○ | ● | ● | Low-Code | Configurable Plotly dashboards |
| Querybook | ○ | ● | ● | SQL Notebook | Collaborative SQL notebooks |
| WrenAI | ○ | ● | ● | BI/Text-to-SQL | Natural language queries |
| Appsmith | ◐ | ● | ● | Low-Code App | Internal data tools |
| Budibase | ◐ | ● | ● | Low-Code App | Fast internal apps |
| DataEase | ○ | ● | ● | No-Code BI | Open source BI dashboards |
| ToolJet | ◐ | ● | ● | Low-Code App | Internal apps/dashboards |
| Ackee | ● | ● | ◐ | Web Analytics | Privacy-respecting analytics |
| GoatCounter | ● | ● | ◐ | Web Analytics | Lightweight, cookieless |
| Open Web Analytics | ● | ● | ● | Web Analytics | Self-hosted alternative |
| OpenReplay | ● | ● | ◐ | Session Replay | Replay and UX debugging |
| Umami | ● | ● | ◐ | Web Analytics | Simple, privacy-first |
| Rybbit | ● | ● | ◐ | Web Analytics | Modern open source analytics |
Analysis Leaders: (excluding new additions)
- Streaming: Grafana (monitoring), Kibana (logs), Dash, Streamlit
- Micro-Batching: Superset, Metabase, Grafana
- Batching: Superset, Metabase, Matplotlib
FUNCTION: GOVERNANCE
Orchestration, quality, security, compliance
Governance × Flow Matrix
| Solution | Stream | Micro-Batch | Batch | Category | Main strength |
|---|---|---|---|---|---|
| Apache Airflow | ○ | ◐ | ● | Orchestration | Batch leader, DAG |
| Dagster | ○ | ◐ | ● | Orchestration | Asset-based, data-aware |
| Prefect | ○ | ◐ | ● | Orchestration | Dynamic workflows |
| Kestra | ◐ | ● | ● | Orchestration | Event-driven, versatile |
| DataHub | ○ | ◐ | ● | Data Catalog | Metadata discovery |
| OpenMetadata | ○ | ◐ | ● | Data Catalog | Open standard, lineage |
| Great Expectations | ○ | ● | ● | Data Quality | Validation, profiling |
| SQLFluff | ○ | ○ | ● | SQL Linter | Code quality |
| Pandera | ○ | ● | ● | Data Quality | Python dataframe validation |
| Evidently | ○ | ● | ● | ML Monitoring | ML drift and quality |
| OpenLineage | ◐ | ● | ● | Data Lineage | Open lineage standard |
| CKAN | ○ | ◐ | ● | Open Data Portal | Open data catalog |
| Apache Ranger | ◐ | ● | ● | Security | Centralized access policies |
| Apache Egeria | ○ | ◐ | ● | Metadata Governance | Cross-enterprise metadata |
| Node-RED | ● | ● | ◐ | Flow Automation | Visual IoT automation |
| Activepieces | ◐ | ● | ● | Automation | No-code workflow automation |
| Huginn | ◐ | ● | ● | Automation | Self-hosted agents and alerts |
| Kepler | ○ | ● | ● | Energy Monitoring | Container energy consumption |
| Scaphandre | ● | ● | ◐ | Energy Monitoring | Real-time energy metrics |
| Flyte | ○ | ◐ | ● | Orchestration | Scalable ML/data workflows |
| Temporal | ◐ | ● | ● | Workflow Engine | Reliable durable workflows |
| Kedro | ○ | ◐ | ● | Pipeline Framework | Reproducible data pipelines |
| Metaflow | ○ | ◐ | ● | ML Pipeline | Netflix ML workflows |
| ZenML | ○ | ◐ | ● | MLOps | Portable ML pipelines |
| Windmill | ◐ | ● | ● | Workflow/Automation | Dev scripts and workflows |
| Hamilton | ○ | ◐ | ● | Pipeline Framework | Dataflow function DAG |
| Astronomer Cosmos | ○ | ◐ | ● | Orchestration | dbt inside Airflow |
| Microsoft Presidio | ○ | ● | ● | PII/Privacy | PII detection/anonymization |
| Faker | ○ | ◐ | ● | Synthetic Data | Fake data generation |
| OPA | ● | ● | ● | Policy Engine | Policies as code (Rego) |
| OpenFGA | ● | ● | ● | Authorization | Fine-grained authorization (ReBAC) |
| Cerbos | ● | ● | ● | Authorization | Decoupled authorization |
| SDV | ○ | ◐ | ● | Synthetic Data | Realistic synthetic data |
| Permify | ● | ● | ● | Authorization | Google Zanzibar authorization |
Governance Leaders: (excluding new additions)
- Streaming: Kestra (event triggers)
- Micro-Batching: Kestra, Great Expectations
- Batching: Airflow (standard), Dagster, DataHub
Recommendation Matrices
Matrix 1: By Use Case × Function
| Use Case | Collection | Transport | Storage | Processing | Analysis | Governance |
|---|---|---|---|---|---|---|
| ** Real-Time Alerts** | Debezium | Kafka | Hudi | Flink | Grafana | Kestra |
| ** Real-Time Dashboard** | Vector | Kafka | ClickHouse | Flink | Kibana | - |
| ** Fraud Detection** | CDC | Kafka | Druid | Flink | Dash | - |
| ** Near-Real-Time BI** | NiFi | Pulsar | Delta Lake | Spark | Superset | Airflow |
| ** Data Warehouse** | Airbyte | - | Iceberg | dbt | Superset | Airflow |
| ** ML Pipeline** | dlt | - | Parquet | Spark | Jupyter | Airflow |
| ** IoT Platform** | MQTT | NATS | Paimon | Flink | Grafana | Kestra |
| ** Web Analytics** | Snowplow | Kafka | ClickHouse | Druid | Posthog | - |
Matrix 2: Full Stack by Flow × All
PURE STREAMING STACK (Latency < 1s)
| Function | Recommended Solution | Alternative |
|---|---|---|
| Collection | Debezium (CDC) | Maxwell, Fluent Bit |
| Transport | Apache Kafka | NATS, Pulsar |
| Storage | Hudi | Paimon, ClickHouse |
| Processing | Apache Flink | Storm, Materialize |
| Analysis | Grafana | Kibana, Dash |
| Governance | Kestra | - |
Architecture Example:
MySQL → Debezium → Kafka → Flink → Hudi → Trino → Grafana
↓
ClickHouse → Superset
MICRO-BATCHING STACK (Latency 1s-5min)
| Function | Recommended Solution | Alternative |
|---|---|---|
| Collection | Apache NiFi | Camel, Bento |
| Transport | Apache Kafka | Pulsar |
| Storage | Delta Lake | Iceberg |
| Processing | Apache Spark | Beam, DuckDB |
| Analysis | Superset | Metabase |
| Governance | Airflow | Dagster |
Architecture Example:
APIs → NiFi → Kafka → Spark Streaming → Delta Lake → Trino → Superset
↓
StarRocks → Grafana (refresh 1min)
BATCHING STACK (Latency > 5min)
| Function | Recommended Solution | Alternative |
|---|---|---|
| Collection | Airbyte | Meltano, dlt |
| Transport | - | NiFi (if needed) |
| Storage | Iceberg | Parquet, Delta Lake |
| Processing | dbt + Spark | Trino, Hop |
| Analysis | Superset | Metabase, Lightdash |
| Governance | Airflow | Dagster, Prefect |
Architecture Example:
Sources → Airbyte → MinIO/S3 (Raw) → Spark → Iceberg (Processed)
↓
dbt → Iceberg (Curated) → Trino → Superset
↓
Airflow (orchestration)
Summary Tables
Top 10 Multi-Function Solutions
| Solution | Functions Covered | Flow Modes | Versatility Score |
|---|---|---|---|
| Apache NiFi | Collection + Transport + Governance | ●●● | |
| Apache Kafka | Transport + Processing | ●●◐ | |
| Apache Spark | Processing + Storage | ◐●● | |
| ClickHouse | Storage + Processing | ●●● | |
| Grafana | Analysis + Monitoring | ●●● | |
| Apache Flink | Streaming processing | ●●◐ | |
| Hudi | Versatile storage | ●●● | |
| Superset | BI analysis | ○●● | |
| Airflow | Batch orchestration | ○◐● | |
| dbt core | SQL transform | ○◐● |
Functional Coverage by Flow
| Flow | Collection | Transport | Storage | Processing | Analysis | Governance | Total |
|---|---|---|---|---|---|---|---|
| ** Streaming** | 33 | 19 | 20 | 26 | 38 | 9 | 145 |
| ** Micro-Batch** | 24 | 20 | 34 | 45 | 59 | 22 | 204 |
| ** Batching** | 12 | 10 | 42 | 35 | 60 | 32 | 191 |
Insight: Micro-batching offers the best overall coverage (204 tool-function combinations).
Gap Analysis by Function
| Function | Streaming Coverage | Batch Coverage | Main Gap |
|---|---|---|---|
| Collection | [OK] Excellent | [OK] Excellent | - |
| Transport | [OK] Excellent | [!] Limited | Batch does not require transport |
| Storage | [!] Limited | [OK] Excellent | Limited streaming formats |
| Processing | [OK] Good | [OK] Excellent | - |
| Analysis | [OK] Good | [OK] Excellent | - |
| Governance | [X] Weak | [OK] Excellent | Missing streaming orchestration |
Recommendations:
- Streaming: Needs orchestration (Kestra in development)
- Batching: Needs real-time transport (not critical)
Two-Dimensional Selection Guide
Step 1: Identify the Function
Question: What is the position in the pipeline?
- Start (Source) → Collection
- Middle (Routing) → Transport
- Base (Data) → Storage
- Middle (Transform) → Processing
- End (Insights) → Analysis
- Cross-cutting (Ops) → Governance
Step 2: Identify the Flow
Question: What is the maximum acceptable latency?
- < 1 second → Streaming
- 1 sec - 5 min → Micro-Batching
- > 5 minutes → Batching
Step 3: Consult the Matrix
Refer to the corresponding section (e.g., “Collection × Streaming”)
Decision Example
Need: E-commerce dashboards with real-time metrics (sales, inventory)
- Function: Analysis (end of pipeline)
- Required latency: 10 seconds → Micro-Batching
- Analysis × Micro-Batch Matrix:
- Options: Superset ●, Grafana ●, Metabase ●
- Choice: Superset (rich BI) + Grafana (monitoring)
Two-Dimensional Architecture Patterns
Pattern 1: Full Lambda (All Functions)
COLLECTION
├─ Streaming: Debezium (CDC)
└─ Batching: Airbyte (Bulk)
TRANSPORT
└─ Streaming: Kafka
STORAGE
├─ Hot: ClickHouse (streaming)
└─ Cold: Iceberg (batch)
PROCESSING
├─ Streaming: Flink
└─ Batching: Spark + dbt
ANALYSIS
├─ Real-time: Grafana
└─ BI: Superset
GOVERNANCE
└─ Batching: Airflow
Pattern 2: Simplified Kappa (Streaming Only)
COLLECTION: Fluent Bit
TRANSPORT: NATS
STORAGE: Paimon
PROCESSING: Flink
ANALYSIS: Grafana
GOVERNANCE: Kestra (event-driven)
Pattern 3: Modern Data Stack (Batch + BI)
COLLECTION: Airbyte
STORAGE: Iceberg on MinIO
PROCESSING: dbt
ANALYSIS: Superset
GOVERNANCE: Airflow + DataHub
Heatmap: Solution Maturity
Maturity Legend
- ● Mature (production-ready, large adoption)
- ◐ Growing (production possible, moderate adoption)
- ○ Emerging (early-stage, limited adoption)
| Collection | Transport | Storage | Processing | Analysis | Governance | |
|---|---|---|---|---|---|---|
| ** Stream** | ● | ● | ◐ | ● | ● | ○ |
| ** Micro** | ● | ● | ● | ● | ● | ◐ |
| ** Batch** | ● | ● | ● | ● | ● | ● |
Insights:
- Mature zone: Batch (all functions) and Micro-Batch
- In-development zone: Streaming Governance, Streaming Storage
- Solid zone: Transport and Processing (all flows)
Cross-References
- Classification by data flow
- Ingestion and Transport
- Storage
- Query and Processing
- Analysis and Output
- Platform Management
Classification Methodology
Function Criteria
Each tool is classified according to its primary function:
- Collection: Initial capture from sources
- Transport: Reliable delivery between systems
- Storage: Persistence and organization
- Processing: Transformation and enrichment
- Analysis: Visualization and BI
- Governance: Orchestration and quality
Flow Criteria
Classification according to 3 modes based on latency observed in production:
- ● Primary: Optimal and recommended usage
- ◐ Secondary: Supported but not optimal
- ○ Not supported: Not recommended
Validation
Each classification is based on:
- Official documentation
- Production use cases
- Community benchmarks
- Established architecture patterns
Document created on: 2025-12-09 Last updated: 2026-06-25 Version: 2.1 Status: [OK] Complete - 263 solutions × 6 functions × 3 flows