OSS Data Tools Landscape

A landscape of open-source solutions for building a composable data platform: an inventory by function, a classification by data flow, licenses, and an interactive report.

Open the interactive report — search, filters (category, section, license, criteria) and sorting (stars, dates), with top/flop and recent/stale cues. Hosted on GitHub Pages, and also works when opened directly.

Catalogues — open-source tools inventoried by function:

Cross-cutting analyses — the same tools, seen differently:

Resources:

Project goals

Market study

In today’s data-driven world, a myriad of open-source solutions are available to process, analyze and manage data. This abundance offers flexibility and power, but can also be overwhelming for organizations trying to pick the right tools for their needs.

The challenge of choice

The open-source community has developed numerous solutions covering the whole data chain (ingestion, storage, query & processing, analysis, platform management).

Overview

While this diversity is a testament to the field’s innovation, it makes decision-making complex and time-consuming for teams building or evolving their data infrastructure.

Our approach: tailored filters for informed decisions

To help with the choice, we propose a set of filters that take into account factors such as:

  1. Company size and scale of data operations
  2. Industry-specific requirements and regulations
  3. Existing technology stack and integration needs
  4. Performance and scalability needs
  5. Level of in-house expertise and resources
  6. Long-term maintainability and community support

By applying these filters, the vast catalogue is narrowed down to a manageable selection of high-value tools for each use case.

Benefits

A working prototype

An application to explore Steam statistics.

It uses Kestra, dbt, Evidence and PostgreSQL.

Get the project: github.com/olexya/data-games-viz