Skip to content

Fix lineage through LATERAL FLATTEN, reused derived aliases, predicate subqueries and star options - #75

Closed
IL-William wants to merge 8 commits into
pondpilot:masterfrom
IL-William:fix/relation-scopes
Closed

IL-William wants to merge 8 commits into
pondpilot:masterfrom
IL-William:fix/relation-scopes

Conversation

@IL-William

Copy link
Copy Markdown

Builds on #74, whose three commits come first here until it lands: the tests
below use the expected_sources helper it adds. The last five commits, four
fixes and their changelog entry, are this PR's.

Summary

Four ways a relation or a scope was misread, each giving wrong column lineage
without raising an issue.

  • LATERAL FLATTEN was not matched by the lineage visitor, so its alias was
    taken for a table name: in
    SELECT f.value::STRING AS tag FROM orders o, LATERAL FLATTEN(input => o.tags) f,
    tag came from a column value of a table F that does not exist, and
    orders.tags was never read.
  • Two derived tables sharing an alias in one statement shared one node. A
    generated point in time query writes one (SELECT DISTINCT k, ts FROM sat_x) AS v per satellite, in separate CTEs, and every consumer then received every
    satellite's lineage.
  • A subquery a predicate reads was analysed with no target, so its select
    list fell back to whatever the enclosing select targets: SELECT k FROM t WHERE x >= (SELECT MAX(y) AS floor_y FROM u) returned k and floor_y, an
    EXISTS (SELECT 1 ...) returned col_0, and inside a CTE the column joined
    the CTE's list, which a star over it then returned. A derived table without an
    alias had no node at all, and the outer select found no table in scope.
  • * EXCLUDE, * EXCEPT and * RENAME were ignored in wildcard expansion.
    With the common idiom SELECT * EXCLUDE (b), UPPER(b) AS b FROM p, the star
    copy of b and the computed b shared one output node, the copy came first,
    and the UPPER(b) derivation was lost.

Changes

One commit per fix, each with its tests, after those of #74:

  • fix(core): read LATERAL FLATTEN as a relation with its own columns
    registers FLATTEN under its alias as a relation with the six columns
    Snowflake documents; VALUE and THIS derive from the input, SEQ, KEY, PATH and
    INDEX read nothing. Each FLATTEN gets a node keyed by its scope, since CTEs
    commonly reuse one alias. Any other LATERAL <function>(...) now raises
    UNSUPPORTED_SYNTAX. The snowflake_lateral_flatten snapshot changes
    accordingly.
  • fix(core): give each derived table its own node when an alias repeats
    counts derived tables per alias and statement: the first keeps the key it
    always had, so statements that do not repeat an alias are unchanged, and each
    later one is keyed alias#n.
  • fix(core): keep a predicate's subquery and an unaliased derived table out of the outputs analyses a subquery read by a predicate, a join condition or a
    grouping expression against a node of its own, labelled (subquery), and
    takes what it projects back from the projection buffer. A scalar subquery in
    the select list is not visited there and is returned as before. A derived
    table without an alias gets a node like an aliased one, registered in scope
    under (derived). This one builds on the previous commit's per occurrence key.
  • fix(core): honour EXCLUDE, EXCEPT and RENAME in a wildcard's expansion
    passes the wildcard's options to expand_wildcard: excluded columns are not
    expanded, renamed ones are expanded under their new name, read from the old
    one. REPLACE and ILIKE are still not read, and a plain SELECT *, f(x) AS x
    is left as it was.

A CTE name defined again in a nested WITH merges the same way the derived
tables did, but CTEs are also looked up by name after the statement has been
walked, so that takes more than a key and is not attempted here.

Each commit message has the details and examples. There is no API or schema
change.

Validation

  • cargo test --workspace --locked, cargo clippy --workspace --locked -- -D warnings, cargo fmt --all -- --check and the schema guard pass on this
    branch.
  • The tests of the last three commits fail without their fix, except the one
    pinning that a select list scalar subquery is still returned, which passes
    either way by design.
  • Found while using flowscope-core as the engine of a dbt column lineage tool.
    On a large Snowflake dbt project: 105 of 111 FLATTEN value pairs had been
    missing, 8 point in time models had their satellites crossed, 33 models
    gained a column their SQL does not return, and 16 edges ran through columns an
    EXCLUDE had removed.

The committed browser WASM is not rebuilt here, since CI builds its own.

🤖 Generated with Claude Code

IL-William and others added 8 commits September 29, 2026 09:48
The parser builds a chain of binary operators left-deep: `a || b || c`
is `(a || b) || c`. The expression walkers recursed into both operands
and counted every operator as one level of nesting, so a flat chain of
more than about 100 operators passed MAX_RECURSION_DEPTH. The walkers
then gave up at the guard, which sits at the leftmost end of the chain,
and the statement was reported as APPROXIMATE_LINEAGE.

Generated surrogate keys reach that length quickly, because every field
is wrapped and joined with a separator:

    SELECT MD5(
        COALESCE(CAST(field_000 AS VARCHAR), '') || '-' ||
        COALESCE(CAST(field_001 AS VARCHAR), '') || '-' ||
        ...
        COALESCE(CAST(field_119 AS VARCHAR), '')
    ) AS row_key
    FROM events

Over 120 fields, row_key was derived from the last 49 only: field_000
to field_070 were missing from its lineage. A key over 50 fields
already loses its first operands.

Walk the left spine of a BinaryOp chain in a loop and visit its operands
left to right, each one level below the chain. The length of a chain no
longer counts as depth, while nesting in right operands, function
arguments and parentheses still does, so the guard keeps protecting the
stack. The four walkers that share MAX_RECURSION_DEPTH use it:
visit_expression_for_subqueries, collect_column_refs,
find_aggregate_function and collect_simple_identifiers. Operands are
still visited in source order.

Tests: the 120-field key above keeps every field and raises no
APPROXIMATE_LINEAGE; unit tests pin the operand order and check that
nesting on the right still trips the guard.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
sqlparser gives TRIM, SUBSTRING, POSITION, CEIL and FLOOR, AT TIME ZONE,
IS [NOT] DISTINCT FROM, the `col:path` and `col[i]` accessors, COLLATE,
OVERLAY, SIMILAR TO, RLIKE, ANY and ALL, and a few more, their own `Expr`
variants rather than function calls. collect_column_refs matched none of
them and ended in a catch-all arm, so a column inside one of these forms
was never a source, and no issue said so:

    SELECT o.id, TRIM(c.name) AS customer_name
    FROM orders o
    JOIN customers c ON o.customer_id = c.id

customer_name came out with no source column. In
`SUBSTR(code, 1, 2) || suffix` only suffix was read, and
`WHERE TRIM(c.status) = 'active'` was not attached to customers.

collect_simple_identifiers, which drives lateral column alias
resolution, already walked SUBSTRING, CEIL and POSITION but not TRIM, so
a key hashed from earlier aliases of the same SELECT list lost every
source:

    SELECT
        MD5(UPPER(TRIM(CAST(order_id AS VARCHAR)))) AS order_key,
        MD5(UPPER(TRIM(CAST(customer_id AS VARCHAR)))) AS customer_key,
        MD5(CONCAT(UPPER(TRIM(CAST(order_key AS VARCHAR))), '||',
                   UPPER(TRIM(CAST(customer_key AS VARCHAR))))) AS link_key
    FROM orders

Make the match in collect_column_refs exhaustive. Every form descends
into the operands it evaluates in the enclosing scope, including named
arguments written `name => value`, the tested value of
`x IN (SELECT ...)`, subscripts and bracket keys. Field names in a path
step and lambda parameters are not read, and subqueries keep their own
scope. The parser keeps `o.items[1]` as the root `o` followed by the
steps `.items` and `[1]`, so the leading dot steps are folded back into
the name: the column read is `o.items`, as without the subscript, and
not a column `o` of orders. With no catch-all arm left, a variant added
to sqlparser is a compile error rather than a column dropped from
lineage.
collect_simple_identifiers walks the same forms, so an alias hidden in
one of them is replaced by its sources instead of being resolved as a
column of a table in scope.

The postgres_array_slicing snapshot changes accordingly: `a[:]`,
`b[:1]`, `c[2:]` and `d[2:3]` now derive from the columns a, b, c and d
rather than from the table node.

Tests: the queries above, SUBSTR, CEIL, POSITION and `:` operands in
Snowflake, the left operand of IN (subquery), a DuckDB lambda and a
subscript of a qualified column in Postgres at the analysis level; a
table of dedicated forms pins what collect_column_refs reads, and checks
that extract_simple_identifiers sees the same bare names.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
sqlparser gives `LATERAL FLATTEN(...)` its own table factor,
`TableFactor::Function`, which the lineage visitor did not match. The
alias was never registered and no issue was raised, so a column read
through it lost its source:

    SELECT o.order_id, f.value::STRING AS tag
    FROM orders o,
    LATERAL FLATTEN(input => o.tags) f

`f` was taken for a table name: tag came out derived from a column
`value` of a table `F` that does not exist, the implied schema listed
`F` beside ORDERS, and orders.tags was read neither for lineage nor for
the implied schema. An unqualified `value` was reported as ambiguous
across the tables in scope, and a FLATTEN over the value of another had
nothing to follow.

Register FLATTEN under its alias, or under its name when it has none,
as a relation with the columns Snowflake documents: SEQ, KEY, PATH,
INDEX, VALUE and THIS. VALUE and THIS carry the element, so they derive
from the columns of the `INPUT =>` argument, or of the first positional
one, with the call as their expression. SEQ, KEY, PATH and INDEX say
where an element sits and read nothing. They are added without the
relation-level edge a source-less projection gets, which would have made
`'n' || f.index` depend on every base relation of the FROM clause. Each
FLATTEN gets its own node, keyed by its scope, since CTEs commonly reuse
one alias such as `f`. A star over the scope now includes the six
columns.

The alias is also recorded as a subquery alias, as a derived table's
is, so that what is read through it stays out of the implied schema,
and the qualified columns of the input are recorded there as projected
columns are: the query above now implies ORDERS(order_id, tags) and
nothing else. With column lineage disabled only that alias is recorded:
at table level the input is a column of a relation already in the FROM
clause, so a FLATTEN node would have no edge, and no column is emitted.

Those six columns are all a FLATTEN has, so the scope records it as a
relation with fixed columns. A bare name it lacks, with no schema to
place it, still resolves to the one other relation of the FROM clause,
as it did while FLATTEN went unregistered: `id` in `SELECT id FROM
orders o, LATERAL FLATTEN(input => o.tags) f` stays orders.id instead of
turning ambiguous, and in `FROM customers, LATERAL FLATTEN(input =>
addresses) a, LATERAL FLATTEN(input => phones) p` the second input is
read from customers. A warning for a name that is ambiguous between
tables no longer lists the FLATTEN beside them. VALUE and THIS read the
same input, so an input that cannot be resolved is reported once.

Any other `LATERAL <function>(...)` now raises UNSUPPORTED_SYNTAX, as
`TABLE(<function>(...))` already does, instead of being dropped
silently.

The snowflake_lateral_flatten snapshot changes accordingly: `value AS
p_id` now derives from b.cool_ids through the FLATTEN node, B's implied
columns gain cool_ids, and the ambiguity warning for `value` is gone.

Tests: the query above with a decoy `value` column on orders, its
implied schema without a schema given, chained flattens back to the
first input, the position columns reading nothing under a CTE joined to
another table, a literal input, one alias reused by two CTEs with named
and positional input, a star over FLATTEN, a bare column and a bare
filter beside a FLATTEN with no schema, a bare input to a second
FLATTEN, an ambiguous input reported once, SPLIT_TO_TABLE reported as
unsupported, and no FLATTEN node or column with column lineage
disabled.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
A derived table's node was keyed by statement and alias. Two derived
tables of one statement sharing an alias in different scopes, as a
generated point in time query writes one `(...) AS v` per satellite,
therefore shared one node and one set of column nodes, and each consumer
received every producer's lineage:

    WITH a_intervals AS (
        SELECT k, ts FROM (SELECT DISTINCT k, ts FROM sat_a) AS v
    ),
    b_intervals AS (
        SELECT k, ts FROM (SELECT DISTINCT k, ts FROM sat_b) AS v
    )
    SELECT a.ts AS a_ts, b.ts AS b_ts
    FROM a_intervals a JOIN b_intervals b ON a.k = b.k

a_ts came out derived from sat_a.ts and sat_b.ts both, and so did b_ts,
with no issue raised.

Count the derived tables of each alias per statement. The first keeps
the key it always had, so nothing changes for a statement that does not
repeat an alias, and each later one is keyed `alias#n`, getting a node
and column nodes of its own. The node keeps the alias as its label. Two
derived tables of one alias in the same scope are invalid SQL, and
lookups already go through the scope, so a per statement count is
enough.

A CTE name defined again in a nested WITH merges the same way, but CTEs
are also looked up by name once the statement has been walked, so that
takes more than a key and is left alone here.

Tests: the query above, each output from its own satellite and two
nodes labelled v; the same alias over twelve satellites, each output
from its own. Both fail without the change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… out of the outputs

A subquery a predicate reads was analysed with no target node, so its
select list fell back to whatever the enclosing select targets: the
statement's output at the top, and inside a CTE that CTE's column list,
which a star over it then returned.

    SELECT k FROM t WHERE x >= (SELECT MAX(y) AS floor_y FROM u)

returned k and floor_y. The same held for IN, for EXISTS, whose
`SELECT 1` came back as col_0, and for a subquery in HAVING, a join
condition or a grouping expression. No issue said so.

Such a subquery is now analysed against a node of its own, labelled
`(subquery)` and keyed per occurrence, which owns its columns and their
lineage, and what it projects is taken back from the projection buffer
once it has been read. A scalar subquery in the select list is not
visited there and is returned as before.

A derived table without an alias, which Snowflake and others accept,
had no node at all: its projection hung on the enclosing target in the
same way, and the outer select found no table in scope and reported so.
It now gets a node like an aliased one, registered in its scope under a
name no query can write, `(derived)`, so that a star over it expands and
a bare column resolves.

Tests: a scalar comparison, IN, a correlated NOT EXISTS, HAVING and a
join condition each return only k; inside a CTE the subquery's column
stays on its own node and out of the star; a select list scalar
subquery is still returned; an unaliased derived table reads like an
aliased one, with no issue. All but the select list case fail without
the change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
expand_wildcard emitted every column of the relation whatever options
the wildcard carried: `SELECT * EXCLUDE (b) FROM p` returned b, and
`SELECT * RENAME (b AS bee) FROM p` returned b under its old name.

The common idiom of replacing a column fared worse:

    SELECT * EXCLUDE (b), UPPER(b) AS b FROM p

The star copy of b and the computed b shared one output node, the copy
came first, and the UPPER(b) derivation was merged away: b came out as
a plain copy of p.b.

The wildcard's options now reach the expansion. A column named by
EXCLUDE, or by EXCEPT where the dialect writes it so, is not expanded; a
column named by RENAME is expanded under its new name, still read from
the old one. Names are compared as the expanded columns' own names are.
REPLACE and ILIKE are still not read. A plain `SELECT *, f(x) AS x`,
which does write two columns named x, is left as it was.

Tests: EXCLUDE with and without parentheses, qualified and inside a CTE;
EXCEPT; RENAME with and without parentheses, read from the old column;
the idiom above, whose b carries UPPER(b). All fail without the change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@IL-William IL-William closed this Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant