Skip to content

Dimensions and relations#

A dimension is an axis of the model, such as snapshot or generator. Declarations are indexed by it, and sum reduces over it.

A relation is a named table between dimensions: a generator's bus, a snapshot's period, or the buses a generator may connect to.

dimensions#

dimensions:
  snapshot: { dtype: int }
  generator: { dtype: str }

Every dimension named anywhere in the file is declared here.

Field
dtype float, int, str, datetime default str
description free text, never parsed default null

A declaration says that the axis exists and what type its labels have. It never lists the labels. The generators, buses and snapshots arrive with the data, as a table, not as a list somebody keeps in step by hand.

Where the members come from#

The engine that binds the data follows three rules, and every engine follows the same three. So two engines given the same file and the same tables build the same model.

  1. The members come from the key named after the dimension. An engine reads generator from the generator table, and from nowhere else. It reads capacity for its values, never for its list of generators, and it does not treat gen_bus as the list either. If a declaration uses generator and no generator table arrives, the engine raises an error that names generator. It does not build an empty axis, because an empty axis would silently delete every row indexed by it. A declared dimension that no declaration uses needs no table.
  2. The members keep the order the table gives them. The engine does not sort them, whether they are strings, integers or dates. shift, sum_back and position() all count along this order, so an engine that sorted snapshot would give shift(dispatch, along=snapshot, offset=1) a different meaning. To get a particular order, write the table in that order.
  3. A table has each coordinate at most once. Two rows for snapshot == 3 is an error that names 3. The engine does not keep the last, keep the first, or add them. At most once, not exactly once: a coordinate with no row is how a parameter declared coverage: masked masks, and an error at the bind for one declared total. Which of the two it is, the declaration says and the table does not (coverage). A relation's table obeys the same rule.

Every dimension has one list of members, and every parameter is lined up against it when the data binds. So if load has 8760 snapshots and price has 8759, the engine raises an error rather than build a model with one snapshot dropped.

relations#

A relation is what makes topology data: a generator sits on a bus, a line has two endpoints, a snapshot falls in a period, and no adjacency matrix or hand-written join appears anywhere. A relation is a table with one column per dimension it relates, and key: is the claim that makes it a map: one row per key tuple, so the other columns are a function of the key. The declaration fixes no direction; the operator that walks the table says which column it consumes and which it produces.

dimensions:
  bus: { dtype: str }
  generator: { dtype: str }
  line: { dtype: str }
  snapshot: { dtype: int }
  period: { dtype: int }
relations:
  gen_bus: { columns: [generator, bus], key: generator } # each generator on one bus
  line_from: { columns: [line, bus], key: line } # two relations onto one dimension
  line_to: { columns: [line, bus], key: line }
  period_of: { columns: [snapshot, period], key: snapshot }
  connection: { columns: [generator, bus] } # no key: a generator may connect to several buses
Field
over required — the columns: a list of dimensions, or a mapping of column name to dimension where two columns share one (roles)
into not a field: a relation declares no direction
key the columns a row is identified by, one name or a list; omitted, the table is a bare relation (below) default none
coverage total, masked — whether every coordinate of the key carries a row default total
description free text, never parsed default null

Every column is over a declared dimension, and its values are checked against that dimension's labels once data is bound — the check that makes sum(by=) safe, and the reason a label set the model only ever selects on is declared as a dimension all the same: nothing is indexed by period above, and where: "period_of == 1" (where strings) is how a declaration selects on it. A relation has at least two columns; a label on one dimension is a parameter over it. A column named like a dimension is over that dimension, so columns: {bus: line} is refused.

The key is the claim#

key: names the columns that are unique together. key: generator says the generator column holds each label once: the table has one row per generator, so the other column is a function of it. key: [generator, period] says the pair holds each combination once. Neither column need be unique on its own: a generator appears once per period, and a period once per generator. The claim is checked at bind: a generator on two buses is refused, where a 0/1 membership parameter would have said so legally and silently (#161). The columns the key determines are the relation's value columns. A key has one column per dimension: it is read at its dimensions, and no frame carries a dimension twice, so key: [bus0, bus1] is refused where both are over bus.

Each cardinality is one declaration, and the key is the side that is one:

to say write checked at bind
many-to-one, each generator on one bus {columns: [generator, bus], key: generator} one row per generator
one-to-many, a bus and its generators the same table: sum(p, by=gen_bus) collects a bus's generators, at(price, by=gen_bus) reads a generator's bus the same
many-to-many, a generator on several buses {columns: [generator, bus]}, no key nothing: a row exists, or it does not
one-to-one not a claim the language has: a key is one set of columns, so the other side stays many

The key is also what decides which walks the table admits:

the walk needs because
sum(x, by=l, over=a, into=b) the key not wholly inside the columns the operand fixes — the into columns and the columns joined on a sum adds its rows up; walked to the key it finds one per coordinate, which is a read
at(x, by=l, over=a, into=b) a key inside the columns the operand fixes — the into columns and the columns joined on a read is one value per coordinate, or it is not a read
shift, sum_back, position a key column over the dimension walked a coordinate is in one group, or it has no neighbour
where: "l == 'north'" a key, and the column compared a value column a comparison is one value per coordinate
where: l (bare) nothing a row exists, or it does not

A bare relation — no key: — is walked by sum alone, with both ends named, and tested by a bare where. That is what a many-to-many relation can say, and all it can say.

coverage says whether a gap is meant#

A partial relation is legal, and coverage: is where the file says it was meant. A key coordinate the table leaves out belongs to no group, so a generator can sit on no bus and a line can have one open end. sum(by=) places such a coordinate's terms nowhere. That is a deliberate shape and a wiring mistake in equal measure, and the two are identical in the data, so the declaration says which:

relations:
  gen_bus: { columns: [generator, bus], key: generator } # total: every generator is on a bus
  line_to: { columns: [line, bus], key: line, coverage: masked } # an open end is meant

The default is total, so a component library that declares its coupling map total turns a port nobody wired from a term that quietly vanishes into an error that names it. "Left out" is spelled by omission: a key coordinate with no row. A value that names no label of its column is an error.

A total relation must carry a row for every coordinate of its key:. Where there is no key, the claim is over every combination of its columns' dimensions.

A walk names its ends#

Every operator that takes by= walks the table between two of its columns: over= the column consumed, into= the column produced, and every other key column joined on — the operand carries its dimension and the result keeps it. A value column not walked is not read: ends below, walked from line to bus1, joins on nothing. A bare relation's columns are all key, so all of them but the two walked are joined on.

dimensions:
  generator: { dtype: str }
  zone: { dtype: str }
  period: { dtype: int }
relations:
  zone_of: { columns: [generator, period, zone], key: [generator, period] } # a generator's zone, per period
parameters:
  demand: { dims: [zone, period] }
  price: { dims: [zone, period] }
variables:
  p: { dims: [generator, period] }
constraints:
  zone_balance: # p[generator, period] → [zone, period]
    dims: [zone, period]
    expression: sum(p, by=zone_of, over=generator, into=zone) >= demand
  history: # p[generator, period] → [generator, zone]: the same table, walked from its other key column
    dims: [generator, zone]
    expression: sum(p, by=zone_of, over=period, into=zone) <= 100
  capped_revenue: # price[zone, period] → [generator, period]: the price of the zone this generator sat in that period
    dims: [generator, period]
    expression: at(price, by=zone_of, over=zone, into=generator) * p <= 1000

What the declaration decides, the call may leave unsaid. Where the key has one column and the key determines one column, the walk is the arrow the key draws, and sum(p, by=gen_bus) and at(price, by=gen_bus) are complete: sum consumes the key and produces the value, at consumes the value and produces the key. Where a side has several candidates — two key columns, two value columns — the call names it, and the refusal lists the candidates. zone_of above has two key columns, so sum names over=, while into=zone could have been left out.

A partition walks a key column and groups by the value columns. shift(x, along=d, by=l), sum_back(x, along=d, by=l) and position(d, by=l) take the one key column over d; the other key columns are joined on, and the group is the value tuple. within= names the value columns the group is made of where the table has several: shift(x, along=snapshot, by=cal, within=week) walks within weeks of a calendar declared once over [snapshot, day, week], and a value column not named is not read.

The rules, each decided at load with a refusal naming the rewrite:

  • over= and into= name columns of the relation by= names, one each or a list each, and no column on both sides. into= is refused without a by=, since a column needs the table that holds it. over= without one names a dimension of the operand instead, which is sum(p, over=period). sum(p, by=gen_bt, into=[bus, technology]) lands one table with two value columns on the product bus × technology in one join; sum(p, by=zone_of, over=[generator, period]) consumes both key columns at once, which is sum(sum(p, by=zone_of, over=generator), over=period) said once; at(tech_cap, by=gen_bt, over=[bus, technology]) reads a two-column slot at each generator.
  • The operand carries every joined column's dimension, each once. The map is read at the key columns not walked, so there is no reading it at a coordinate that lacks them; two joined columns over one dimension have nothing to tell apart.
  • A produced dimension the operand already carries is joined on too. sum(load * p, by=gen_bus) with load[snapshot, bus] restricts each term to the row where the generator's bus is the row's bus — a masked sum, which is what the join says.
  • at reads one value. Its key lies inside into= and the joined columns, or the call is refused; a bare relation is never read by at.
  • sum adds its rows up. So the reverse holds: a sum whose key lies inside into= and the joined columns finds one row per coordinate and adds up nothing, which is a read — it is refused toward at. sum walks to a value column; at walks to the key.
  • A partition walks the one key column over the dimension it walks, and groups by the value columns within= names — all of them where it names none. within= naming a key column is refused, and a bare relation partitions nothing. The group may hold two columns over one dimension, a pair of buses say: a partition lands nothing, so nothing needs the dimension twice.
  • A by= list walks each relation by its declared arrow. by=[a, b] is one grouping, so no column keyword has anything to name; every relation in it consumes the same dimension, joins on its own other columns, and no two produce the same dimension.
  • A where comparison reads a value column of a keyed relation at its key. zone_of == 'north' reads the one value column; ends.bus0 != ends.bus1 names the columns where there are several. The frame carries the key's dimensions, and two relations compared have keys over the same dimensions and columns over one. A bare name — where: gen_bus — tests that a row exists: at the key for a keyed relation, at every column for a bare relation.
  • Every column is over a declared dimension, every column name is distinct, the key names columns the relation has, and does not name all of them.

Every relation name joins the flat namespace, so a relation may not shadow a dimension. generator's map onto bus is gen_bus, never a second bus.

Roles#

A list under columns: names each column after its dimension. Two columns over one dimension need names of their own, and the mapping form gives them:

relations:
  ends: { columns: { line: line, bus0: bus, bus1: bus }, key: line } # a line's two ends, one table
  rep_of: { columns: { snapshot: snapshot, rep: snapshot }, key: snapshot } # the representative snapshot

sum(f, by=ends, over=line, into=bus1) - sum(f, by=ends, over=line, into=bus0) is the nodal balance through one table where two relations did it before, and where: "ends.bus0 != ends.bus1" excludes a self-loop by comparing two of its columns.

rep_of relates a dimension to itself, which is how a clustered year is run on a few typical days: every snapshot names the one that stands for it. Nothing changes in the rules — snapshot is consumed and rep produced, both over one dimension, so the frame is unchanged through sum(by=) and at(by=) alike:

constraints:
  representative: # every snapshot takes its representative's value
    dims: [snapshot]
    expression: p == at(p, by=rep_of)
  weighted: # the snapshots a representative stands for, summed onto it
    dims: [snapshot]
    expression: sum(p, by=rep_of) <= 100

A self-map is directional exactly as far as its key says. key: snapshot makes rep a function of snapshot, so the arrow runs from a snapshot to its representative: at reads along it and sum collects against it, the inverse of a many-to-one map being one-to-many, reachable as a grouping and never as a function. Two steps along the arrow are two nested calls. Without a key the same two columns are an undirected relation — a neighbour table — which sum walks either way and nothing reads. Selecting the representatives themselves, the rows where the map is the identity, is not a comparison the language has, since a relation is never compared to a dimension; declare a bool parameter for them.

How the map is supplied#

gen_bus is a source key like any other, carrying one column per column declared, named after the column:

sources = {
    'generator': ['g1', 'g2', 'g3'],
    'gen_bus': pl.DataFrame({'generator': ['g1', 'g2'], 'bus': ['north', 'south']}),
}

A partial map is the rows it has. g3 is in no row, so g3 sits on no bus — absence is the absent row, exactly as it is for a parameter, and a null in any column is refused for saying both at once. A keyed table holds one row per key tuple, and a value matching no label of its column's dimension is a typo rather than a new member. Values are never inferred from the parameters that use a dimension: inferring would let a mistyped label extend the label set instead of being rejected.

Supplying it this way touches no table but its own, which is what a caller who did not generate the index needs: a model can be extended with a relation the same way it can be extended with a parameter. A column of a dimension's index named after a relation is refused rather than read — an index may carry any other extra, and this one would be a map read by accident.

Dimension, relation or parameter?#

Every column of data is one of the three. What decides which is what the math does with the column, not what the column holds:

The column… is declared as because
is an axis: something is indexed by it, or an aggregation lands terms on it a dimension its members are the coordinate set every table over it is reindexed onto
has one value per member of a dimension, or per tuple of several — a generator's bus, a line's two ends, a generator's zone by period a relation with that key it is a map every operator walks, and its values are checked against the dimensions they name
relates members of two dimensions many-to-many, with nothing to weigh — which buses a generator may connect to a relation with no key sum walks it with both ends named, and a bare where tests it. Nothing reads it, because there is no one value to read
relates members of two dimensions many-to-many, with a weight per pair — a link's efficiency to each bus, a cycle's lines a parameter over both the weight is the data, its row set is the relation, and the aggregation is sum(w * x, over=a)
is a label set the model only selects on or counts within — a period, a season, a zone a dimension, and a keyed relation onto it the membership check is worth one line and one member list
scales terms — a coefficient, a bound, an offset a parameter (float or int) arithmetic is over numbers (dtype)
is a per-row attribute the math only selects on — a fuel, a constraint's sense a str parameter it names rows rather than scaling them, and no set is declared to check its values against
is a mask a bool parameter a bare name in a where is its own answer

Two rules follow from the table. If b has one value per a, then b is a relation keyed by a, and not a dimension: a dims product over two dimensions that depend on each other, cut back with a mask, is the shape that relations replaces.

And everything under dimensions: is an axis. A dimension is never legal where a value belongs, because it is a coordinate space and not data. To use a dimension's coordinates as data, declare a parameter over it. python -m math_spec check advises on a declared dimension that nothing is indexed by, nothing aggregates into and no relation has a column over (errors).