Reference

The reference library

201 entries, grouped by the layer of the stack each one belongs to. Reading a layer top to bottom is the fastest way to understand that part of the architecture.

Foundations

The architectural ideas the rest of the stack assumes: what a lakehouse is, and how it differs from a lake or a warehouse.

17 entries
Bronze LayerIn the Medallion Architecture that structures modern data lakehouses, data does not simply arrive and immediately become analytically ready.Data FabricModern enterprises rarely store their data in a single place.Data GravityData Gravity is an analogy coined by Dave McCrory in 2010 to describe the phenomenon by which large concentrations of data attract applications, services.Data LakeThe explosion of digital information over the last decade created a massive storage problem.Data LakehouseA comprehensive definition of Data Lakehouse architecture, combining data warehouse reliability with data lake scalability via open table formats.Data MeshFor most of the 2010s, the standard blueprint for a modern data platform involved building a centralized data lake, staffing a central data engineering team.Data SwampA comprehensive guide to what causes a Data Lake to become a Data Swamp, how to recognize the warning signs, and the governance practices that prevent it.Data WarehouseA data warehouse is a centralized repository engineered specifically to store highly structured, historical data.Gold LayerAt the top of the Medallion Architecture sits the Gold Layer: the final destination for data that has been ingested, cleaned, validated, enriched, and now.Kappa ArchitectureUnderstanding Kappa Architecture, the simplified alternative to Lambda that treats everything as a stream.Lambda ArchitectureLambda Architecture is a data deployment model introduced by Nathan Marz designed to handle massive quantities of data by taking advantage of both batch and.Medallion ArchitectureWhen organizations first started building data lakehouses, they faced a structuring problem.Open LakehouseAn open lakehouse is a data architecture that stores data in open file and table formats on commodity object storage, so that any compliant engine can read and write it without vendor lock-in.Polyglot PersistenceFor a long time, the default answer to any data storage question in enterprise software development was a relational database.Separation of Compute and StorageSeparation of Compute and Storage is an architectural principle in which the processing layer (compute engines that execute queries and transformations) and.Silver LayerIn the Medallion Architecture, data quality is not enforced at the point of ingestion.Zero-ETLData pipelines are expensive to build and expensive to maintain.

Storage format

How bytes are laid out on object storage, and how they are encoded and compressed for analytical reads.

26 entries
Amazon S3Amazon Simple Storage Service (Amazon S3) is an object storage service provided by Amazon Web Services (AWS) that offers industry-leading scalability, data.Avro FormatWhile Apache Parquet dominates the field of analytical data storage, it is not the only file format found in a modern data lakehouse.Azure Blob StorageAzure Blob Storage is Microsoft's massively scalable object storage solution for the cloud.Bloom FiltersA Bloom Filter is a space-efficient probabilistic data structure invented by Burton Howard Bloom in 1970.Column-Level StatisticsColumn-Level Statistics are metadata measurements computed over the values in individual columns of data files, stored alongside the data to enable two.Columnar FormatsColumnar formats are data storage layouts where data is physically organized and stored by column, rather than by row.Data FileIf you peel back the complex layers of an open data lakehouse (past the query engines, past the Catalog, past the metadata tree of Manifest Lists and.Dictionary EncodingDictionary Encoding is a highly effective data compression technique predominantly used in columnar storage formats like Apache Parquet and Apache ORC.File Block SizeFile Block Size (also referred to as row group size or split size) is a critical physical configuration parameter in big data storage systems.File FormatIn the architecture of a modern data lakehouse, there is a strict separation between the Table Format (like Apache Iceberg or Delta Lake) and the File Format.Google Cloud Storage (GCS)Google Cloud Storage (GCS) is a managed, highly scalable object storage service provided by Google Cloud Platform (GCP).GZIP CompressionGZIP (GNU zip) is one of the most widely used lossless data compression utilities in the history of computing.LZ4 CompressionLZ4 is a lossless data compression algorithm focused on incredibly fast compression and decompression speeds.Min-Max StatisticsMin-Max Statistics are per-column metadata stored alongside data in Parquet files and in Iceberg Manifest Files, recording the minimum value, maximum value.MinIOMinIO is an open-source, high-performance, distributed object storage server.Object StorageObject Storage is a data storage architecture designed to manage massive amounts of unstructured and structured data.ORC FormatApache ORC (Optimized Row Columnar) is the second major columnar file format found in modern data lakehouses.Parquet FormatApache Parquet is an open-source, columnar file format designed specifically for fast data processing and massive storage efficiency in the Hadoop and data.Row-Oriented FormatsRow-oriented formats are data storage layouts where the data associated with a single record (a row) is stored contiguously on physical storage media.Run-Length Encoding (RLE)Understanding Run-Length Encoding (RLE), a foundational compression algorithm for sorted columnar data.S3 API CompatibilityS3 API Compatibility refers to the industry-wide phenomenon where competing cloud providers, independent software vendors, and hardware manufacturers have.Small File ProblemThe Small File Problem refers to the performance and operational degradation that occurs when a data lake or lakehouse accumulates a very large number of.Snappy CompressionSnappy is a fast, lossless data compression and decompression library written in C++ and originally developed by Google.Storage LayerIn the architecture of a modern data lakehouse, the Storage Layer is the foundational bedrock upon which everything else is built.Target File SizeTarget File Size is the configured desired size (in bytes) for data files written to or produced by a data lakehouse table format like Apache Iceberg.Zstandard (Zstd)Zstandard, commonly abbreviated as Zstd, is a fast, lossless data compression algorithm developed by Yann Collet at Facebook (Meta).

Table format

The metadata that turns a directory of files into a table with transactions, schema evolution, and time travel.

50 entries
ACID TransactionsACID is the foundational set of properties that define the correctness guarantees a data storage system must provide for transactions to be considered.Apache HudiApache Hudi (Hadoop Upserts Deletes and Incrementals) is the third major pillar of the Open Table Format ecosystem alongside Apache Iceberg and Delta Lake.Apache IcebergAs data lakes grew in popularity, organizations quickly realized that simply dumping raw Parquet or ORC files into Amazon S3 was not a viable long-term.Apache PaimonApache Paimon is the youngest and most architecturally distinctive of the major Open Table Formats.Apache XTable (OneTable)Apache XTable, originally released by Onehouse as "OneTable" and donated to the Apache Software Foundation, represents a fundamentally different category of.Branching (WAP)Software engineering relies heavily on version control systems like Git.Commit (Iceberg)In a traditional file-system-based data lake, "committing" data usually just meant finishing the upload of a Parquet file to a directory.CompactionA modern data lakehouse is often fed by continuous, real-time data streams (like Apache Kafka or Flink).Copy-on-Write (CoW)In a traditional relational database, if you execute an UPDATE statement to change a user's address, the database engine physically overwrites the specific.Delete FilesIn an immutable storage layer like Amazon S3, you cannot open a Parquet file, locate a specific row, and hit the "delete" key.Delta LakeDelta Lake is one of the three dominant Open Table Formats that define the modern data lakehouse.Delta UniFormDelta UniForm, short for Universal Format, is a feature introduced in Delta Lake 3.0 that allows a Delta Lake table to be simultaneously readable as an.Equality DeletesWhen dealing with high-velocity streaming data, such as a Change Data Capture (CDC) pipeline ingesting thousands of database updates per second, data.Expire SnapshotsA defining feature of Apache Iceberg is its ability to create a new Snapshot for every single transaction.Format ConversionFormat Conversion is the process of reading data stored in one physical file format and rewriting it in a different physical file format.Format InteroperabilityThe promise of the Open Data Lakehouse is simple to state and brutally difficult to execute: store data once, on cheap cloud object storage, and make it.Hidden PartitioningPartitioning is a core optimization strategy in massive data lakehouses.Manifest FileIn the Apache Iceberg metadata hierarchy, the Manifest File is the critical layer sitting directly above the raw data.Manifest ListIn the hierarchical metadata tree of Apache Iceberg, the Snapshot defines the state of the table, but it does not directly list the millions of data files.Merge-on-Read (MoR)While Copy-on-Write (CoW) provides blistering read performance, its massive Write Amplification makes it unsuitable for high-frequency updates, such as.Metadata LayerIf data files (like Parquet or ORC) are the muscle of a data lakehouse, and compute engines (like Spark or Dremio) are the brain, the Metadata Layer is the.Metadata LogApache Iceberg achieves ACID transactions on object storage without requiring a continuous compute engine.Metadata PointerIn an open data lakehouse architecture based on Apache Iceberg, a single table might consist of millions of Parquet data files, thousands of Avro manifest.Metadata TranslationMetadata Translation is the technical process of reading the metadata representation of a data table in one Open Table Format and generating an equivalent.Open Table FormatsA definitive, detailed look guide into Open Table Formats, exploring the architectural shift in approach that bridges the gap between data lakes and data warehouses, featuring an exhaustive analysis of Apache Iceberg, Delta Lake, and Apache Hudi.Optimistic Concurrency Control (OCC)When multiple systems attempt to write data to the same table at the exact same time, a database must have a mechanism to resolve the conflict.Partition EvolutionIn the lifecycle of a data lakehouse, data volume rarely remains static.Partition SpecPartitioning is a fundamental technique in data engineering.Position DeletesWithin the Merge-on-Read (MoR) architecture of Apache Iceberg, there are two methods for creating a logical tombstone to hide a record.Read AmplificationRead Amplification is the phenomenon where a query must read more data from storage than is logically required to satisfy the query's result, due to the.Remove Orphan FilesData lakehouse operations are inherently distributed and prone to environmental failures.Rewrite Data FilesIn Apache Iceberg, the abstract concept of "Compaction" is practically executed using a specific maintenance API called rewriteDataFiles.Rewrite ManifestsWhile rewriteDataFiles focuses on optimizing the physical Parquet data, it does not optimize the metadata layer.RollbackDespite the best data quality checks and Write-Audit-Publish patterns, human error inevitably occurs in data engineering.Schema EvolutionA comprehensive guide to Schema Evolution in Apache Iceberg, detailing how metadata-only operations provide safe, instantaneous updates to data structures.Schema SpecA comprehensive guide to the Schema Spec in Apache Iceberg, detailing how strict column ID tracking enables safe, instantaneous schema evolution without rewriting data.Sequence NumberIn a highly concurrent data lakehouse, determining the exact order in which events occurred is critical.SnapshotIn a traditional relational database, if you execute an UPDATE statement to change a user's address, the database physically overwrites the old address on.Snapshot IsolationWhen multiple systems interact with the same data simultaneously, chaos can easily ensue.Sort Order SpecWhile the Partition Spec determines how data is logically divided into coarse-grained directories (like by year or month), the Sort Order Spec determines how.Staged CommitsThe Write-Audit-Publish (WAP) pattern is the gold standard for maintaining data quality in a lakehouse.Strict MetricsWhen a query engine like Apache Spark or Trino runs a query against a data lakehouse, its primary goal during the planning phase is to read as little data as.Table FormatIf you look inside a traditional data lake, you will find a collection of directories and files.Table MaintenanceTable Maintenance is the collection of recurring operational procedures that keep Apache Iceberg tables performant, storage-efficient, and operationally.Table UUIDWhen a user interacts with a database, they use human-readable names. They write SELECT * FROM sales_data.Tagging (Iceberg)While Time Travel allows you to query historical data by providing a specific Snapshot ID or Timestamp, remembering a random 19-digit number like.Time TravelIn traditional relational databases, querying the past is extremely difficult.Transaction LogThe Transaction Log is one of the most fundamental data structures in computer science.Write AmplificationWrite Amplification is a phenomenon in data storage and processing systems where the actual amount of data written to storage is significantly larger than.Write-Audit-Publish (WAP)The "silent failure" is the most dangerous event in data engineering. A pipeline succeeds, no errors are thrown, and data is written to production.

Catalog

The service that tracks which tables exist, where their metadata lives, and who is allowed to touch them.

17 entries
Attribute-Based Access Control (ABAC)Attribute-Based Access Control is an access management model that makes authorization decisions based on the attributes of the subject (who is requesting.AWS Glue Data CatalogThe AWS Glue Data Catalog is Amazon Web Services' fully managed, serverless metadata catalog service: the central metadata registry for the entire AWS.Catalog MigrationCatalog Migration is the process of transitioning Apache Iceberg tables from one catalog implementation to another, from Hive Metastore to a REST Catalog.Credential VendingCredential Vending is the security mechanism in the Apache Iceberg REST Catalog specification that enables catalog services to dynamically generate.Dremio ArcticA historical overview of Dremio Arctic, the Git-for-data catalog that evolved into Apache Polaris and the Nessie open-source project.Dynamic CatalogsDynamic Catalogs refers to the architectural pattern of configuring query engines and data platforms to connect with multiple Iceberg catalog instances.Fine-Grained Access Control (FGAC)Fine-Grained Access Control refers to the collection of data access enforcement mechanisms that operate below the table level: controlling which rows.Hadoop CatalogThe Hadoop Catalog (also referred to as the Filesystem Catalog or HadoopCatalog in the Iceberg codebase) is the simplest possible catalog implementation for.Hive Metastore (HMS)The Hive Metastore is the metadata management service at the heart of the Apache Hive data warehouse ecosystem, and by extension the foundational catalog.Iceberg CatalogThe Iceberg Catalog is the central nervous system of any Apache Iceberg deployment.JDBC CatalogThe JDBC Catalog is a catalog implementation for Apache Iceberg that uses any JDBC-compatible relational database as its persistent metadata backend.Polaris CatalogApache Polaris is the premier open-source implementation of the Apache Iceberg REST Catalog specification: a production-grade, vendor-neutral catalog service.Project NessieProject Nessie is an open-source transactional catalog for Apache Iceberg that introduces Git-like version control semantics to data lakehouse metadata.REST CatalogThe Iceberg REST Catalog specification, commonly called the REST Catalog or IRC, is the single most architecturally significant standard to emerge from the.Role-Based Access Control (RBAC)Role-Based Access Control is the dominant access management model in enterprise data infrastructure: a system that grants permissions to named roles, and.TabularTabular was a managed data lakehouse platform built entirely around Apache Iceberg, founded in 2021 by the original creators of the Apache Iceberg project.Unity CatalogUnity Catalog is Databricks' enterprise governance layer for the lakehouse: a centralized metadata, access control, and data discovery service that manages.

Compute

The engines that plan and execute queries, and the optimisations that make them fast over object storage.

43 entries
Amazon AthenaAmazon Athena is a serverless, interactive query service provided by Amazon Web Services (AWS) that allows users to analyze data directly in Amazon Simple.Apache DorisApache Doris is a modern, open-source Massively Parallel Processing (MPP) analytical database designed for blazing-fast, real-time data warehousing.Apache FlinkApache Flink is an open-source, unified stream-processing and batch-processing framework.Apache SparkApache Spark is a unified analytics engine designed for large-scale data processing.Broadcast JoinA Broadcast Join is a distributed join algorithm where the smaller of the two join inputs is completely replicated ("broadcast") to every worker node in the.CachingCaching in data systems is the practice of storing copies of frequently accessed data in a faster, closer storage layer so that subsequent requests for the.ClickHouseClickHouse is an open-source, column-oriented database management system (DBMS) built expressly for online analytical processing (OLAP).Compute EngineIn traditional data warehousing, storage and compute were inextricably linked.Cost-Based Optimizer (CBO)A Cost-Based Optimizer (CBO) is the intelligent "brain" within a modern relational database or distributed compute engine.Data SkewData Skew in distributed query execution is the unequal distribution of data or computational work across the worker nodes of a cluster, where some workers.Data SkippingData Skipping is a query optimization technique in which a query engine uses pre-computed metadata statistics about data files to determine, at planning.DatabricksAn extensive overview of Databricks, the unified data analytics platform that pioneered the data lakehouse paradigm and developed Delta Lake and Apache Spark.DeserializationAn in-depth look at deserialization and its performance impacts on analytical query engines.Distributed ComputeA foundational overview of distributed compute architectures in data processing, explaining master-worker topologies, data shuffling, and fault tolerance.DremioDremio is an open lakehouse platform purpose-built for Apache Iceberg, providing a high-performance query engine, a semantic layer, an integrated Apache.DuckDBDuckDB is an open-source, in-process SQL OLAP database management system.File SkippingFile Skipping is the collective term for the multi-level system of techniques that allow a data lakehouse query engine to identify and eliminate data files.Google BigQueryGoogle BigQuery is a fully managed, serverless, enterprise data warehouse that enables scalable analysis over petabytes of data.Hash JoinA Hash Join is a join algorithm that uses a hash table to match rows from two input relations based on their join key values.Hilbert CurvesThe Hilbert Curve is a mathematically elegant space-filling curve first described by mathematician David Hilbert in 1891.Indexing (Data Lakes)Indexing in the context of data lakes and lakehouses refers to the set of auxiliary data structures that enable query engines to quickly locate relevant data.Join StrategiesA Join Strategy is the specific algorithm a query engine uses to physically implement a SQL JOIN operation: matching rows from two (or more) tables based on.Materialized ViewsA Materialized View is a pre-computed, physically stored result of a SQL query whose output is saved to storage and can be queried directly, rather than.MPP (Massively Parallel Processing)Massively Parallel Processing (MPP) is an architectural design paradigm for distributed computing and databases.Out-of-Memory (OOM) ErrorsAn Out-of-Memory (OOM) error occurs when a query execution process attempts to allocate more memory than is available to it, either on a single node.Partition PruningPartition Pruning is the most powerful form of data skipping available to query engines in the data lakehouse.Predicate PushdownPredicate Pushdown is a query optimization technique in which filter conditions (predicates) from the WHERE clause of a SQL query are applied as early as.PrestoPresto (often referred to as PrestoDB to distinguish it from its fork, Trino) is an open-source, distributed SQL query engine designed for running.Projection PushdownProjection Pushdown (also called column pruning) is a query optimization technique in which only the specific columns required by a SQL query are read from.Pushdown OptimizationPushdown Optimization (specifically Predicate Pushdown and Projection Pushdown) is one of the most critical performance techniques in distributed data.Query ExecutionQuery Execution is the runtime phase in which a database or query engine carries out the physical execution plan produced by the query planner, reading data.Query PlanningQuery Planning is the process by which a database or query engine transforms a declarative SQL query into a detailed, optimized execution plan: a concrete.Rule-Based Optimizer (RBO)A Rule-Based Optimizer (RBO) is a critical component of a database's query planning phase.SerializationSerialization is the process of translating data structures or object state into a format that can be stored (for example, in a file or memory buffer) or.ShuffleIn distributed query execution, a Shuffle (also called an Exchange or Repartition) is the operation of redistributing data rows across the worker nodes of a.SnowflakeSnowflake is a fully managed cloud data platform that fundamentally reshaped the data warehousing industry.Sort-Merge JoinA Sort-Merge Join (SMJ) is a join algorithm that sorts both input relations by their join key and then performs a single linear merge pass through both.Spilling to DiskSpilling to Disk is a query engine mechanism that writes intermediate query results (hash tables, sort buffers, shuffle data) to local disk storage when the.SQL DialectsSQL (Structured Query Language) is the lingua franca of data analytics.StarRocksStarRocks is an open-source, next-generation Massive Parallel Processing (MPP) database designed for blazing-fast, real-time analytics.TrinoTrino (formerly known as PrestoSQL) is a highly parallel and distributed open-source SQL query engine.Vectorized ExecutionA detailed explanation of vectorized execution, the hardware-optimized processing model that allows modern compute engines to achieve blistering speeds.Z-OrderingZ-Ordering is a multi-dimensional data clustering technique used in data lakehouse environments to physically co-locate records with similar values across.

Interchange

The in-memory format and wire protocols that move results between systems without paying serialisation costs.

3 entries

Movement

How data arrives and keeps arriving: ingestion, transformation, orchestration, and streaming.

15 entries
Apache AirflowApache Airflow is an open-source platform created by Airbnb in 2014 and later donated to the Apache Software Foundation.Batch ProcessingBatch Processing is the execution of a series of jobs in a computer program without manual intervention (non-interactive).Change Data Capture (CDC)Change Data Capture (CDC) is a set of software design patterns and technologies used to determine and track the data that has changed within a source.DagsterDagster is an open-source data orchestration platform designed to address some of the architectural limitations of Apache Airflow.Data PipelineA Data Pipeline is an automated set of processes and infrastructure that extracts data from various source systems, transforms it into a clean and usable.dbt (data build tool)Understanding dbt, the transformative framework that brought software engineering best practices to SQL-based data transformations.Directed Acyclic Graph (DAG)A Directed Acyclic Graph (DAG) is a conceptual mathematical model heavily utilized in computer science, specifically within the area of data engineering and.ELT (Extract, Load, Transform)For decades, the dominant pattern for moving data from source systems into analytical databases was ETL: Extract, Transform, Load.ETL (Extract, Transform, Load)Before a business analyst can run a query against a data warehouse, someone has to get the data there.Eventual ConsistencyIn any discussion of data reliability in distributed systems, the conversation inevitably polarizes around two competing consistency models: ACID (Atomicity.Micro-batchingExploring Micro-batching, the architectural compromise that simulates streaming using rapid, tiny batch jobs.OrchestrationIn data engineering, Orchestration is the automated configuration, coordination, and management of complex computer systems, software, and services.PrefectExploring Prefect, the dynamic, Python-native workflow orchestration framework.Streaming DataStreaming Data refers to data that is continuously generated by thousands of data sources, which typically send in the data records simultaneously, and in.Strong ConsistencyStrong Consistency is the most demanding correctness guarantee a distributed storage system can provide.

Semantics

The layer that maps physical tables to business meaning, so queries are written against concepts rather than schemas.

12 entries
Data LineageData Lineage is the documented record of how data moves, transforms, and evolves through an analytical system, from its origination at source systems through.Data ModelingData Modeling is the process of creating a visual and logical representation of either a whole information system or parts of it to communicate connections.Data QualityData Quality is the set of properties, practices, and enforcement mechanisms that ensure data assets in a lakehouse are fit for the purposes for which they.Dimension TableUnderstanding Dimension Tables, the descriptive context that gives meaning to analytical data.Dimensional ModelingDimensional Modeling is a specialized data design methodology primarily utilized for data warehouses, data marts, and modern data lakehouses.Fact TableIn dimensional modeling and data warehousing (specifically within a Star Schema or Snowflake Schema), a Fact Table is the central table that stores the.Knowledge GraphsA Knowledge Graph is a structured representation of knowledge as a network of entities and the relationships between them.OntologyAn ontology is a formal, explicit specification of a shared conceptualization: a structured vocabulary that defines the types of entities that exist in a.Semantic LayerOne of the most persistent and costly problems in enterprise analytics is metric inconsistency.Slowly Changing Dimensions (SCD)Slowly Changing Dimensions (SCD) is a fundamental concept in data warehousing that deals with a critical problem: How do you handle dimensional data that.Snowflake SchemaThe Snowflake Schema is a logical arrangement of tables in a multidimensional database that is an extension and variation of the Star Schema.Star SchemaUnderstanding the Star Schema, the fundamental dimensional modeling technique optimized for analytical query performance.

Agents & AI

How language models and agents consume governed lakehouse data, and what they need from the layers beneath.

18 entries
Agentic AnalyticsAgentic Analytics represents the next frontier in business intelligence and data engineering.Agentic WorkflowsAn Agentic Workflow is a structured sequence of AI-driven steps in which one or more autonomous agents coordinate to accomplish a complex, multi-stage.AI AgentsAn AI Agent is an autonomous software system that uses a Large Language Model (LLM) as its core reasoning engine to perceive its environment, form plans.Autonomous AnalyticsAutonomous Analytics is the discipline and set of technologies that enable AI systems to independently discover, retrieve, analyze, and interpret enterprise.Context WindowThe context window of a Large Language Model is the total amount of text, measured in tokens, that the model can process and reason about simultaneously.Hallucination MitigationHallucination in Large Language Models refers to the generation of content that is factually incorrect, unverifiable, or entirely fabricated, but presented.Large Language Models (LLMs)A Large Language Model (LLM) is an artificial intelligence model trained on massive quantities of text data to understand, generate, and reason with human.Model Fine-TuningModel fine-tuning is the process of taking a pre-trained Large Language Model and continuing its training on a curated, domain-specific dataset to adapt its.Multi-Agent SystemsA Multi-Agent System (MAS) is an architecture in which multiple independent AI agents collaborate to accomplish complex tasks that exceed what any single.Observability (AI Systems)AI System Observability is the practice of instrumenting, collecting, and analyzing telemetry data from Large Language Model applications and agentic systems.Prompt EngineeringPrompt Engineering is the discipline of designing, structuring, and optimizing the text inputs (prompts) provided to a Large Language Model to elicit.Retrieval-Augmented Generation (RAG)Retrieval-Augmented Generation (RAG) is an AI architecture pattern that enhances a Large Language Model's responses by dynamically retrieving relevant.Semantic SearchSemantic Search is a search methodology that understands the intent and contextual meaning behind a query rather than performing a literal word-for-word.Text EmbeddingsA text embedding is a numerical representation of a piece of text (a word, sentence, paragraph, or entire document) expressed as a dense vector of.Text-to-SQLText-to-SQL is the task of automatically translating a natural language question or instruction into a valid SQL query that, when executed against the target.Tool Use (Function Calling)Tool Use, also called Function Calling, is the capability of modern Large Language Models to generate structured requests to execute predefined external.Vector DatabasesA Vector Database is a specialized database management system engineered to store, index, and efficiently query high-dimensional vector embeddings at scale.Vector SearchVector Search (also called semantic search or similarity search) is a retrieval technique that finds results based on the conceptual meaning and semantic.