Find and hire tech professionals

Dice backs GlossaryTech to keep it free for the community

Data Science

Prophet

A forecasting tool designed to predict time based patterns in data. It is commonly used in Data Science and Machine Learning projects to estimate future values.

Apache Hudi

A data management framework designed for large scale data lakes. It is often used in Big Data platforms with tools like Apache Spark to process changing datasets reliably.

Apache NiFi

A tool that moves and routes data between systems using a visual flow builder. It is often used for Data Pipeline automation in Big Data environments alongside Apache Kafka.

Delta Lake

A storage layer that adds reliability and structure to data lakes. It is often used in data platforms to support machine learning and analytics on large datasets.

MLFlow

A platform that helps manage and track experiments during machine learning development. It allows teams to compare models, parameters, and results in one place.

Model Serving

A way to make trained machine learning models available so applications can use them to make predictions. It focuses on reliability, speed, and scaling models in real-world systems.

Feast

A feature store that helps manage and serve data for machine learning models. It ensures that training and production systems use consistent feature values.

Dagster

A data orchestration platform designed to build reliable data pipelines. It helps teams manage dependencies and data workflows more safely.

dbt

A tool that helps teams transform and model data directly inside a data warehouse. It is widely used in modern analytics workflows to manage data logic as code.

Trino

A distributed SQL query engine designed to query large datasets across multiple data sources. It is commonly used in data platforms for fast interactive analytics.

Serving

A process that makes trained models or services available so applications can use them in real time. It is commonly discussed in machine learning systems.

Data Factory

A cloud service that helps build and manage data pipelines. It is commonly used to move and transform data between systems.

UDF

A custom function created by users to extend system capabilities. It is commonly used in databases and data processing tools.

Prefect

A workflow orchestration tool used to manage data pipelines. It focuses on reliability and task observability.

Looker

A business intelligence platform used to explore and visualize data. It helps teams make data driven decisions.

Tecton

A feature platform that helps manage and serve data for machine learning models. It ensures consistency between training and production.

Iceberg

A table format designed to manage large datasets reliably. It is often used in analytics and machine learning platforms.

Cube.js

A framework that helps build analytics applications on top of databases. It is commonly used to create dashboards using SQL data.

Parquet

A column-oriented file format optimized for storing and querying large datasets. It is widely used in big data and analytics systems.

Shiny

A web framework that allows users to build interactive applications directly from data analysis code. It is commonly used for dashboards.

Polars

A data processing library designed for fast analytics on large datasets. It is commonly used as an alternative to traditional data frames.

GeoPandas

A Python library used to work with geographic data. It extends pandas with spatial operations.

Streamlit

A framework that allows developers to build interactive data applications quickly. It is commonly used in Machine Learning projects.

Anomalo

A data quality platform that detects issues in datasets automatically. It is used to improve trust in analytics and machine learning data.

Tez

A data processing framework designed to build high performance data pipelines. It is commonly used in big data environments.

Sign up for updates
straight to your inbox