Skip to main content

🧑‍🍳 Spice.ai OSS Cookbook

122 guides and samples to help you build data-grounded AI apps and agents with Spice.ai Open-Source. Find ready-to-use examples for data acceleration, AI agents, LLM memory, and more.

Contribute to the Cookbook on GitHub!

Featured Recipes

Federated SQL Query

Join S3 and PostgreSQL data in one SQL query.

Run Llama3 Locally

Use Llama models from HuggingFace with Spice.

Data Acceleration with Cayenne

Speed up queries using Cayenne.

LLM Memory

Persistent memory for language models

Sample Applications and Guides

Example apps and guides for real-world Spice.ai usage and best practices.

Command Query Responsibility Segregation (CQRS)

Sample application implementing the CQRS pattern with Spice.

Intelligent Security Copilot

Use AI to analyze real-time data access patterns and detect potential security risks.

Core Features

Start with the core capabilities: federated SQL query, data acceleration, async queries, hybrid search, and AI inference in SQL.

Federated SQL Query

Join data from S3 and PostgreSQL in a single SQL query, then accelerate both locally.

Cayenne Data Accelerator

Accelerate a local copy of a dataset stored in S3 using Cayenne.

Async Queries

Submit long-running SQL queries and retrieve the results asynchronously.

Hybrid-Search with RRF

Combine full-text and vector search using Reciprocal Rank Fusion (RRF) for improved search results.

AI SQL Function

Invoke LLMs directly within SQL queries using the AI SQL function.

Models, AI, and Agents

Connect to hosted and local AI models, and build intelligent agents using Spice.ai.

AI SQL Function

Invoke LLMs directly within SQL queries using the AI SQL function.

Azure OpenAI Models

Use Azure OpenAI models for vector search and chat over structured and unstructured data.

OpenAI Models

Use OpenAI language and embedding models with Spice.

Running Llama3 Locally

Use the Llama family of models locally from HuggingFace using Spice.

Filesystem Hosted Model

Serve a model stored on the local filesystem.

OpenAI SDK

Use the OpenAI SDK to connect to models hosted on Spice.

OpenAI Responses API

Use the OpenAI Responses API with Spice.

LLM Memory

Persistent memory for language models.

Text to SQL (NSQL)

Ask natural language (NLP) questions of your datasets using the built-in text-to-SQL tool.

Generative Visualizations

Generate SQL queries and interactive charts from natural language questions.

Nvidia NIM on Kubernetes

Deploy Nvidia NIM infrastructure on Kubernetes with GPUs, connected to Spice.

Nvidia NIM on AWS EC2

Deploy Nvidia NIM on a GPU-optimized AWS EC2 instance, connected to Spice.

Searching GitHub Files

Search GitHub files with embeddings and vector similarity search.

xAI Models

Use xAI models such as Grok.

DeepSeek Model

Use DeepSeek model through Spice.

Model Context Protocol (MCP)

Connect to MCP servers and use MCP tools with Spice.

Spice as an MCP Server

Run Spice as an MCP server and connect AI assistants such as Claude Desktop, Cursor, or VS Code.

Amazon S3 Vectors

Use Amazon S3 Vectors to store embeddings and perform efficient vector search.

Data Acceleration, Materialization, and Federation

Optimize query performance with local acceleration, data materialization, and federation techniques.

Cayenne Data Accelerator

Accelerate a local copy of a dataset stored in S3 using Cayenne.

DuckDB Data Accelerator

Accelerate data locally using DuckDB.

PostgreSQL Data Accelerator

Materialize data into an attached PostgreSQL instance. Available in Spice.ai Enterprise.

SQLite Data Accelerator

Accelerate data locally using SQLite.

Apache Arrow Data Accelerator

Accelerate data in memory using Apache Arrow.

Hashed Partitioning with Cayenne

Prune data on categorical columns, such as IDs, using hashed partitioning.

Dataset Partitioning

Partition accelerated datasets so queries skip partitions they do not need.

Accelerated Views

Pre-calculate and materialize derived data for faster queries.

Acceleration Snapshots

Bootstrap accelerations from snapshots in object storage to skip cold starts. Available in Spice.ai Enterprise.

Dual-Dataset Registration

Serve queries immediately while a large table accelerates in the background.

Serializable Transactions

Commit gated, serializable transactions on Cayenne tables and write the results back to PostgreSQL.

Indexes on Accelerated Data

Create and manage indexes on accelerated data.

Change Data Capture (CDC)

Stream inserts, updates, and deletes from source databases to keep accelerated datasets current.

PostgreSQL CDC

Stream changes from PostgreSQL using logical replication, without Debezium or Kafka.

PostgreSQL Catalog CDC

Discover every table in a PostgreSQL database and keep a local copy of each current with CDC.

MySQL CDC

Stream changes from MySQL using the binary log, without Debezium or Kafka.

AWS Aurora MySQL CDC

Stream changes from an AWS Aurora MySQL cluster using the binary log.

MongoDB Change Streams

Stream changes from a MongoDB collection using Change Streams.

DynamoDB Streams

Stream inserts, updates, and deletes from a DynamoDB table using DynamoDB Streams.

Debezium CDC from Postgres

Stream changes from PostgreSQL using Debezium CDC.

Debezium CDC with SASL/SCRAM

Stream MySQL changes using Debezium with SASL/SCRAM authentication.

Search & Embeddings

Search data with full-text, vector, and hybrid search, using embeddings and vector engines.

Hybrid-Search with RRF

Combine full-text and vector search using Reciprocal Rank Fusion (RRF) for improved search results.

Searching GitHub Files

Search GitHub files with embeddings and vector similarity search.

Amazon S3 Vectors

Use Amazon S3 Vectors to store embeddings and perform efficient vector search.

Full-Text Search

Retrieve records matching keywords using BM25 scoring.

Elasticsearch Full-Text and Vector Search

Use Elasticsearch for both full-text and vector search. Available in Spice.ai Enterprise.

Data Connectors

Connect to databases, data warehouses, data lakes, and APIs, and query them with SQL.

PostgreSQL Connector

Connect to and query PostgreSQL databases.

AWS RDS PostgreSQL

Connect to AWS RDS PostgreSQL instances.

Supabase PostgreSQL

Connect to Supabase PostgreSQL databases.

MySQL Connector

Connect to and query MySQL databases.

AWS RDS Aurora MySQL

Connect to AWS RDS Aurora with MySQL compatibility.

PlanetScale MySQL

Connect to PlanetScale MySQL databases.

ClickHouse Connector

Connect to and query ClickHouse databases.

Databricks Connector

Connect to and query Databricks instances using Delta Lake or Spark Connect.

Delta Lake Connector

Query data from Delta Lake tables.

Dremio Connector

Connect to and query a self-hosted Dremio instance running in Docker.

DuckDB Connector

Query DuckDB databases with sample TPCH data.

DynamoDB Connector

Query data from an AWS-hosted DynamoDB table.

Elasticsearch Connector

Query Elasticsearch indices using federated SQL. Available in Spice.ai Enterprise.

File Connector

Query data from local files.

FTP Connector

Query data from FTP and SFTP servers.

GitHub Connector

Connect to and query GitHub data.

GraphQL Connector

Query data from GraphQL endpoints.

HTTP Connector

Query data from HTTP(S) endpoints, such as REST APIs.

MSSQL Connector

Connect to and query across multiple Microsoft SQL Server instances.

ODBC Connector

Connect to databases using ODBC. Available in Spice.ai Enterprise.

Oracle Connector

Connect to and accelerate data from Oracle databases.

Amazon Redshift

Read and write TPC-H data with Amazon Redshift.

Glue Connector

Query tables in an AWS Glue Data Catalog.

S3 Connector

Query data from S3 compatible storage.

ScyllaDB Connector

Query ScyllaDB clusters using federated SQL. Available in Spice.ai Enterprise.

SharePoint Connector

Connect to SharePoint and OneDrive for Business.

SMB Connector

Query data files on SMB network shares.

Snowflake Connector

Connect to and query Snowflake databases.

Snowflake DML

Ingest data from an HTTP API and write it to a Snowflake table with INSERT.

Spice.ai Cloud Connector

Connect to the Spice.ai Cloud Platform.

Apache Spark Connector

Connect to and query Apache Spark.

IMAP Emails

Federated SQL query of mail across IMAP email servers.

MongoDB Connector

Connect to and query MongoDB databases.

Live Orders Analytics with Apache Kafka Data Connector

Combine real-time data streaming from Kafka with other datasets using Spice.

Catalog Connectors

Connect to data catalogs to discover and query their tables.

PostgreSQL Catalog CDC

Discover every table in a PostgreSQL database and keep a local copy of each current with CDC.

Spice.ai Cloud Platform Catalog

Connect to the Spice.ai Cloud Platform catalog.

Databricks Unity Catalog

Connect to Databricks Unity catalog.

Unity Catalog

Connect to an open-source Unity Catalog instance.

Iceberg Catalog Connector

Connect to Iceberg catalog with support for reading and writing Iceberg tables.

Iceberg Hadoop Catalog

Connect to Iceberg Hadoop catalogs, locally or on S3-compatible object storage.

AWS Glue Catalog

Query tables registered in an AWS Glue Data Catalog.

DuckLake Catalog

Discover and query all schemas and tables in a DuckLake catalog.

PostgreSQL Catalog

Discover and query all schemas and tables in a PostgreSQL database.

MySQL Catalog

Discover and query all databases and tables in a MySQL server.

Microsoft SQL Server Catalog

Discover and query all schemas and tables in a SQL Server database.

Visualization

Visualize data with BI and analytics tools.

Generative Visualizations

Generate SQL queries and interactive charts from natural language questions.

Sales BI with Apache Superset

Visualize data in Spice with Apache Superset.

Grafana Datasource

Add Spice as a Grafana datasource.

API Clients

Use API clients for data access and integration.

Python ADBC Client

Query Spice using ADBC and Parameterized Queries with Python.

Java JDBC Client

Query Spice.ai using the Java JDBC client.

Scala JDBC Client

Query Spice.ai using the Scala JDBC client.

cURL with Spice.ai Cloud

Run SQL against Spice.ai Cloud over HTTP using cURL.

Spice CLI with Spice.ai Cloud

Run SQL against Spice.ai Cloud using the Spice CLI.

Deployment

Deploy Spice.ai in different environments.

Deploying to Kubernetes

Deploy Spice.ai on Kubernetes using the Helm chart.

Running in Docker

Run Spice.ai in Docker containers.

Sidecar Deployment Architecture

Run Spice alongside the application on the same host for low-latency access.

Microservice Deployment Architecture

Run Spice as an independent service, optionally with replicas behind a load balancer.

Cloud Connect on a Development Machine

Connect a local Spice instance to Spice Cloud, deploy changes without restarting, and deliver secrets.

Performance and Benchmarking

Measure and optimize performance with benchmarks and best practices for your Spice.ai deployment.

TPC-H Benchmarking

Load TPC-H benchmark data and run the benchmark queries.

SQL Results Caching

Cache query results in memory for faster repeated queries.

Caching Accelerator

Cache HTTP-based datasets with stale-while-revalidate (SWR) support.

Indexes on Accelerated Data

Create and manage indexes on accelerated data.

Configuration

Configure data refresh, retention, scheduling, and data quality for accelerated datasets.

Data Retention Policy

Evict accelerated data older than a specified duration.

Refresh Data Window

Refresh only recent data into accelerated datasets.

Advanced Data Refresh

Advanced configuration for data refresh.

Data Quality with Constraints

Enforce data quality constraints on accelerated datasets.

Cron Dataset Schedules

Schedule dataset refreshes using cron syntax.

SDKs

Use SDKs for different programming languages.

OpenAI SDK

Use the OpenAI SDK to connect to models hosted on Spice.

Rust SDK

Query Spice.ai using the Rust SDK.

Python SDK

Query Spice.ai using the Python SDK.

Go SDK

Query Spice.ai using the Go SDK.

Spice.js JavaScript (Node.js) SDK

Query Spice.ai using the JavaScript (Node.js) SDK with examples.

Java SDK

Query Spice.ai using the Java SDK.

.NET SDK

Query Spice.ai from C# using the .NET SDK.

Security

Secure your Spice.ai deployment and data access with encryption, authentication, and authorization.

TLS Encryption

Enable encryption in transit using TLS.

API Key Authentication

Secure access with API key authentication.

Mutual TLS (mTLS)

Authenticate clients with certificates using mutual TLS. Available in Spice.ai Enterprise.

Authorization

Add multi-tenancy, row-level security, PII masking, and RBAC with Cedar policies. Available in Spice.ai Enterprise.

Advanced Topics

Replicate datasets locally, distribute queries across nodes, and work with JSON data.

Local Dataset Replication

Link datasets in a parent/child relationship within the current Spicepod.

Distributed Query

Run queries across multiple nodes using schedulers and executors.

JSON Strings

Work with JSON strings using JSON functions.