DuckDB Lance Integration
DuckDB's lance core extension
reads Lance tables out of a SeaweedFS table bucket over S3, without the
catalog. That works because a table bucket's layout is a valid Lance dataset
directory — see SeaweedFS Lance Catalog.
Verified against DuckDB 1.5.5 by the integration suite in
test/s3tables/catalog_duckdb_lance/.
Setup
INSTALL lance;
LOAD lance;
CREATE SECRET seaweedfs (
TYPE lance,
ACCESS_KEY_ID '…',
SECRET_ACCESS_KEY '…',
REGION 'us-east-1',
ENDPOINT 'http://localhost:8333',
ALLOW_HTTP true,
VIRTUAL_HOSTED_STYLE_REQUEST false
);
Those are object_store's key names, the same ones the namespace vends in
storage_options.
Reading a table
A table created through the catalog lives at s3://<bucket>/<namespace>/<table>:
SELECT count(*) FROM __lance_scan('s3://vectors/ml/embeddings');
SELECT id, title
FROM __lance_scan('s3://vectors/ml/embeddings')
WHERE id < 5;
SELECT id
FROM lance_vector_search('s3://vectors/ml/embeddings', 'vector',
[1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0], k := 3);
Why __lance_scan and not FROM 's3://…'
DuckDB's replacement scan — the ergonomic SELECT * FROM 's3://…' form —
recognises a Lance dataset by a .lance path suffix. Tables created through
this catalog do not have one, deliberately: the catalog entry is the dataset
directory, a table name may not contain a dot, and a suffix would leak into ARNs
and bucket policies.
-- a dataset written to a .lance path: recognised
SELECT count(*) FROM 's3://vectors/ml/exported.lance';
-- a table created through the catalog: not recognised
SELECT count(*) FROM 's3://vectors/ml/embeddings'; -- Catalog Error
SELECT count(*) FROM __lance_scan('s3://vectors/ml/embeddings'); -- works
Known limitation
ATTACH 's3://bucket/namespace' AS ns (TYPE lance) — the directory-namespace
form in the extension's own documentation — fails in DuckDB 1.5.5 with
Cannot launch in-memory database in read-only mode, whatever options are
passed. Use the scan functions above.
See also
Introduction
- Quick Start with weed mini
- Simplest S3 Bucket and User Setup
- Components
- Blob Store Architecture
- Getting Started
- Production Setup
- A typical step‐by‐step example
- Benchmarks
- FAQ
- Applications
API
Configuration
- Replication
- Store file with a Time To Live
- Failover Master Server
- Erasure coding for warm storage
- EC Bitrot Detection
- Server Startup via Systemd
- Environment Variables
Filer
- Filer Setup
- Directories and Files
- File Operations Quick Reference
- Data Structure for Large Files
- Filer Data Encryption
- Filer Commands and Operations
- Filer JWT Use
- TUS Resumable Uploads
Filer Stores
- Filer Cassandra Setup
- Filer Redis Setup
- Super Large Directories
- Path-Specific Filer Store
- Choosing a Filer Store
- Customize Filer Store
Management
Advanced Filer Configurations
- Migrate to Filer Store
- Add New Filer Store
- Filer Store Replication
- Filer Active Active cross cluster continuous synchronization
- Filer as a Key-Large-Value Store
- Path Specific Configuration
- Filer Change Data Capture
- Filer Operation Serialization
FUSE Mount
- Mount on Windows
- FIO benchmark
- fstab and systemd mount
- POSIX Compliance
- Distributed POSIX Locks
- P2P reading in weed mount
- Mount over the Internet
WebDAV
SFTP Server
Cloud Drive
- Cloud Drive Benefits
- Cloud Drive Architecture
- Configure Remote Storage
- Azure Blob Storage Authentication
- Mount Remote Storage
- Cache Remote Storage
- Cloud Drive Quick Setup
- Gateway to Remote Object Storage
AWS S3 API
- Amazon S3 API
- Supported APIs vs Minio
- S3 Lifecycle
- S3 Lifecycle vs Volume TTL
- S3 Conditional Operations
- S3 CORS
- S3 Object Lock and Retention
- S3 Object Versioning
- S3 RenameObject
- S3 API Benchmark
- S3 API FAQ
- S3 Bucket Quota
- S3 Rate Limiting
- S3 API Audit log
- S3 Nginx Proxy
- Docker Compose for S3
S3 Table Bucket
- S3 Table Bucket
- S3 Table Bucket Commands
- S3 Tables Security
- SeaweedFS Iceberg Catalog
- Iceberg REST Catalog API
- Iceberg Table Maintenance
- SeaweedFS Lance Catalog
- Lance Maintenance Worker
Iceberg Integrations
- Spark Iceberg Integration
- Trino Iceberg Integration
- Dremio Iceberg Integration
- DuckDB Iceberg Integration
- Doris Iceberg Integration
- RisingWave Iceberg Integration
- Lakekeeper Iceberg Integration
Lance Integrations
S3 Authentication & IAM
- S3 Configuration - Start Here
- S3 Credentials (
-s3.config) - OIDC Integration (
-s3.iam.config) - Kubernetes ServiceAccount Authentication (IRSA-style)
- S3 Policy Variables
- S3 Policy Conditions
- S3 Bucket Policies
- Amazon IAM API
- AWS IAM CLI
- weed shell - Shell IAM Commands
Server-Side Encryption
S3 Client Tools
- AWS CLI with SeaweedFS
- s3cmd with SeaweedFS
- rclone with SeaweedFS
- restic with SeaweedFS
- nodejs with Seaweed S3
Machine Learning
HDFS
- Hadoop Compatible File System
- run Spark on SeaweedFS
- run HBase on SeaweedFS
- Run Trino on SeaweedFS
- Hadoop Benchmark
- HDFS via S3 connector
Replication and Backup
- Async Replication to another Filer [Deprecated]
- Async Backup
- Async Filer Metadata Backup
- Async Replication to Cloud [Deprecated]
- Kubernetes Backups and Recovery with K8up
Metadata Change Events
Messaging
- Structured Data Lake with SMQ and SQL
- Seaweed Message Queue
- SQL Queries on Message Queue
- SQL Quick Reference
- PostgreSQL-compatible Server weed db
- Pub-Sub to SMQ to SQL
- Kafka to Kafka Gateway to SMQ to SQL
Use Cases
Operations
- System Metrics
- weed shell
- Data Backup
- Deployment to Kubernetes and Minikube
- Helm Chart Recipes
- Deployment with seaweed-up
Rust Volume Server
Advanced
- Large File Handling
- Optimization
- Optimization for Many Small Buckets
- Volume Management
- Tiered Storage
- Cloud Tier
- Cloud Monitoring
- Load Command Line Options from a file
- SRV Service Discovery
- Volume Files Structure
Security
- Security Overview
- Security Configuration
- Cryptography and FIPS Compliance
- Run Blob Storage on Public Internet