The Iceberg catalog is checked against Spark, Trino, ClickHouse, Doris,
Dremio and RisingWave. The Lance one had only its own reference client,
which is the same thing as checking it against ourselves.
LanceDB connects with connect_namespace("rest", ...), which speaks the
routes this catalog implements, so the suite exercises the protocol rather
than our idea of it: list the catalog, open a table through it, read the
schema, run a vector search and a filtered scan, create a table, and read
the same dataset straight off its URI with no catalog at all.
table_names -> ['lancedb-p0guidmm$ml$embeddings']
open_table -> 64 rows
search -> [1, 0, 2]
create_table -> 4 rows, listed by the catalog
direct read without the catalog -> 64 rows
Seeding is pylance, because the namespace records where a table lives and
does not carry its data. That split is the design rather than a limit of
the test.
One interop note the test encodes: a gateway without STS vends
storage_options carrying an endpoint and a region but no credentials, and
LanceDB uses what the namespace vends on some paths. The container gets
credentials in its environment as well, which is what a deployment without
STS would do.
Claude-Session: https://claude.ai/code/session_01Rkp1Mw5E89Jp6dzJFYiMrm
LanceDB Integration Test
Drives the SeaweedFS Lance Namespace with LanceDB, the way
catalog_spark, catalog_trino and catalog_clickhouse drive the Iceberg REST
catalog with their engines.
Why a real client
Every serious bug in this catalog so far looked correct to a request written by hand: a deregister that deleted the dataset, an S3 door that refused every Lance file, a namespace that listed tables it would then deny. A hand-built HTTP test checks the shape of a response. A client checks whether the response is usable — that the location it hands back, the storage options beside it and the layout rules on the S3 door all line up at once.
LanceDB connects with connect_namespace("rest", ...), which speaks the routes
this catalog implements, so what runs here is the protocol rather than our own
idea of it.
What it does
TestLanceDBNamespace:
- Starts a
weed minicluster with S3 and the Lance Namespace enabled. - Creates a table bucket declared
LANCE, so the catalog refuses tables of any other format in it. - Builds
Dockerfile.client(LanceDB, pylance, lance-namespace) and runslancedb_ops.pyagainst the namespace.
Inside the container:
| Step | What it proves |
|---|---|
| seed a table | the namespace's location and credentials are enough to write |
table_names |
the catalog is browsable through LanceDB |
open_table |
LanceDB resolves a table through the catalog and reads it |
| schema check | the vector column survived the round trip |
search(...) |
ANN search works on data behind SeaweedFS |
where("id < 5") |
so does the scan path, not only the index |
create_table |
reports what the client sees for an operation this catalog does not serve |
direct lance.dataset(uri) |
the catalog stays optional; the dataset opens without it |
The seeding is pylance rather than LanceDB, because the namespace records where a table lives and does not carry its data. That split is the design, not a limitation of the test.
Running it
cd test/s3tables/catalog_lancedb
(cd ../../../weed && go build .) # the harness runs this binary
go test -run TestLanceDBNamespace -v -timeout 30m .
Skipped without Docker, and in -short mode. The first run builds the client
image, which takes a few minutes; later runs reuse it.