Data Science

Delta Lake

A storage layer that adds reliability and structure to data lakes. It is often used in data platforms to support machine learning and analytics on large datasets.

MLFlow

A platform that helps manage and track experiments during machine learning development. It allows teams to compare models, parameters, and results in one place.

Model Serving

A way to make trained machine learning models available so applications can use them to make predictions. It focuses on reliability, speed, and scaling models in real-world systems.

Feast

A feature store that helps manage and serve data for machine learning models. It ensures that training and production systems use consistent feature values.

Dagster

A data orchestration platform designed to build reliable data pipelines. It helps teams manage dependencies and data workflows more safely.

dbt

A tool that helps teams transform and model data directly inside a data warehouse. It is widely used in modern analytics workflows to manage data logic as code.

Trino

A distributed SQL query engine designed to query large datasets across multiple data sources. It is commonly used in data platforms for fast interactive analytics.

Serving

A process that makes trained models or services available so applications can use them in real time. It is commonly discussed in machine learning systems.

Data Factory

A cloud service that helps build and manage data pipelines. It is commonly used to move and transform data between systems.

UDF

A custom function created by users to extend system capabilities. It is commonly used in databases and data processing tools.

Prefect

A workflow orchestration tool used to manage data pipelines. It focuses on reliability and task observability.

Looker

A business intelligence platform used to explore and visualize data. It helps teams make data driven decisions.

Tecton

A feature platform that helps manage and serve data for machine learning models. It ensures consistency between training and production.

Iceberg

A table format designed to manage large datasets reliably. It is often used in analytics and machine learning platforms.

Cube.js

A framework that helps build analytics applications on top of databases. It is commonly used to create dashboards using SQL data.

Parquet

A column-oriented file format optimized for storing and querying large datasets. It is widely used in big data and analytics systems.

Shiny

A web framework that allows users to build interactive applications directly from data analysis code. It is commonly used for dashboards.

Polars

A data processing library designed for fast analytics on large datasets. It is commonly used as an alternative to traditional data frames.

GeoPandas

A Python library used to work with geographic data. It extends pandas with spatial operations.

Streamlit

A framework that allows developers to build interactive data applications quickly. It is commonly used in Machine Learning projects.

Anomalo

A data quality platform that detects issues in datasets automatically. It is used to improve trust in analytics and machine learning data.

Tez

A data processing framework designed to build high performance data pipelines. It is commonly used in big data environments.

ELT

A data integration approach where data is loaded first and transformed later. It is commonly used in modern analytics platforms.

FAISS

A library for efficient similarity search and clustering of dense vectors. It is widely used in machine learning systems.

Prophet

A forecasting tool designed to predict time based patterns in data. It is commonly used in Data Science and Machine Learning projects to estimate future values.

Sign up for updates
straight to your inbox