Dimensions and relations#
A dimension is an axis of the model, such as snapshot or generator.
Declarations are indexed by it, and sum reduces over it.
A relation is a named table between dimensions: a generator's bus, a snapshot's period, or the buses a generator may connect to.
dimensions#
Every dimension named anywhere in the file is declared here.
| Field | ||
|---|---|---|
dtype |
float, int, str, datetime |
default str |
description |
free text, never parsed | default null |
A declaration says that the axis exists and what type its labels have. It never lists the labels. The generators, buses and snapshots arrive with the data, as a table, not as a list somebody keeps in step by hand.
Where the members come from#
The engine that binds the data follows three rules, and every engine follows the same three. So two engines given the same file and the same tables build the same model.
- The members come from the key named after the dimension. An engine reads
generatorfrom thegeneratortable, and from nowhere else. It readscapacityfor its values, never for its list of generators, and it does not treatgen_busas the list either. If a declaration usesgeneratorand nogeneratortable arrives, the engine raises an error that namesgenerator. It does not build an empty axis, because an empty axis would silently delete every row indexed by it. A declared dimension that no declaration uses needs no table. - The members keep the order the table gives them. The engine does not
sort them, whether they are strings, integers or dates.
shift,sum_backandposition()all count along this order, so an engine that sortedsnapshotwould giveshift(dispatch, along=snapshot, offset=1)a different meaning. To get a particular order, write the table in that order. - A table has each coordinate at most once. Two rows for
snapshot == 3is an error that names3. The engine does not keep the last, keep the first, or add them. At most once, not exactly once: a coordinate with no row is how a parameter declaredcoverage: maskedmasks, and an error at the bind for one declaredtotal. Which of the two it is, the declaration says and the table does not (coverage). A relation's table obeys the same rule.
Every dimension has one list of members, and every parameter is lined up against
it when the data binds. So if load has 8760 snapshots and price has 8759, the
engine raises an error rather than build a model with one snapshot dropped.
relations#
A relation is what makes topology data: a generator sits on a bus, a line has
two endpoints, a snapshot falls in a period, and no adjacency matrix or
hand-written join appears anywhere. A relation is a table with one column per
dimension it relates, and key: is the claim that makes it a map: one row per
key tuple, so the other columns are a function of the key. The declaration
fixes no direction; the operator that walks the table says which column it
consumes and which it produces.
dimensions:
bus: { dtype: str }
generator: { dtype: str }
line: { dtype: str }
snapshot: { dtype: int }
period: { dtype: int }
relations:
gen_bus: { columns: [generator, bus], key: generator } # each generator on one bus
line_from: { columns: [line, bus], key: line } # two relations onto one dimension
line_to: { columns: [line, bus], key: line }
period_of: { columns: [snapshot, period], key: snapshot }
connection: { columns: [generator, bus] } # no key: a generator may connect to several buses
| Field | ||
|---|---|---|
over |
required — the columns: a list of dimensions, or a mapping of column name to dimension where two columns share one (roles) | |
into |
not a field: a relation declares no direction | |
key |
the columns a row is identified by, one name or a list; omitted, the table is a bare relation (below) | default none |
coverage |
total, masked — whether every coordinate of the key carries a row |
default total |
description |
free text, never parsed | default null |
Every column is over a declared dimension, and its values are checked against
that dimension's labels once data is bound — the check that makes sum(by=)
safe, and the reason a label set the model only ever selects on is declared
as a dimension all the same: nothing is indexed by period above, and
where: "period_of == 1" (where strings) is
how a declaration selects on it. A relation has at least two columns; a label on
one dimension is a parameter over it. A column named like a dimension is over
that dimension, so columns: {bus: line} is refused.
The key is the claim#
key: names the columns that are unique together. key: generator says the
generator column holds each label once: the table has one row per
generator, so the other column is a function of it. key: [generator, period]
says the pair holds each combination once. Neither column need be unique on its
own: a generator appears once per period, and a period once per generator.
The claim is checked at bind: a generator on two buses is refused, where a
0/1 membership parameter would have said so legally and silently
(#161). The columns
the key determines are the relation's value columns. A key has one column per
dimension: it is read at its dimensions, and no frame carries a dimension
twice, so key: [bus0, bus1] is refused where both are over bus.
Each cardinality is one declaration, and the key is the side that is one:
| to say | write | checked at bind |
|---|---|---|
| many-to-one, each generator on one bus | {columns: [generator, bus], key: generator} |
one row per generator |
| one-to-many, a bus and its generators | the same table: sum(p, by=gen_bus) collects a bus's generators, at(price, by=gen_bus) reads a generator's bus |
the same |
| many-to-many, a generator on several buses | {columns: [generator, bus]}, no key |
nothing: a row exists, or it does not |
| one-to-one | not a claim the language has: a key is one set of columns, so the other side stays many |
The key is also what decides which walks the table admits:
| the walk | needs | because |
|---|---|---|
sum(x, by=l, over=a, into=b) |
the key not wholly inside the columns the operand fixes — the into columns and the columns joined on |
a sum adds its rows up; walked to the key it finds one per coordinate, which is a read |
at(x, by=l, over=a, into=b) |
a key inside the columns the operand fixes — the into columns and the columns joined on |
a read is one value per coordinate, or it is not a read |
shift, sum_back, position |
a key column over the dimension walked | a coordinate is in one group, or it has no neighbour |
where: "l == 'north'" |
a key, and the column compared a value column | a comparison is one value per coordinate |
where: l (bare) |
nothing | a row exists, or it does not |
A bare relation — no key: — is walked by sum alone, with both ends named,
and tested by a bare where. That is what a many-to-many relation can say,
and all it can say.
coverage says whether a gap is meant#
A partial relation is legal, and coverage: is where the file says it was
meant. A key coordinate the table leaves out belongs to no group, so a
generator can sit on no bus and a line can have one open end. sum(by=) places
such a coordinate's terms nowhere. That is a deliberate shape and a wiring
mistake in equal measure, and the two are identical in the data, so the
declaration says which:
relations:
gen_bus: { columns: [generator, bus], key: generator } # total: every generator is on a bus
line_to: { columns: [line, bus], key: line, coverage: masked } # an open end is meant
The default is total, so a component library that declares its coupling map
total turns a port nobody wired from a term that quietly vanishes into an
error that names it. "Left out" is spelled by omission: a key coordinate with
no row. A value that names no label of its column is an error.
A total relation must carry a row for every coordinate of its key:. Where
there is no key, the claim is over every combination of its columns' dimensions.
A walk names its ends#
Every operator that takes by= walks the table between two of its columns:
over= the column consumed, into= the column produced, and every other
key column joined on — the operand carries its dimension and the
result keeps it. A value column not walked is not read: ends below, walked
from line to bus1, joins on nothing. A bare relation's columns are all
key, so all of them but the two walked are joined on.
dimensions:
generator: { dtype: str }
zone: { dtype: str }
period: { dtype: int }
relations:
zone_of: { columns: [generator, period, zone], key: [generator, period] } # a generator's zone, per period
parameters:
demand: { dims: [zone, period] }
price: { dims: [zone, period] }
variables:
p: { dims: [generator, period] }
constraints:
zone_balance: # p[generator, period] → [zone, period]
dims: [zone, period]
expression: sum(p, by=zone_of, over=generator, into=zone) >= demand
history: # p[generator, period] → [generator, zone]: the same table, walked from its other key column
dims: [generator, zone]
expression: sum(p, by=zone_of, over=period, into=zone) <= 100
capped_revenue: # price[zone, period] → [generator, period]: the price of the zone this generator sat in that period
dims: [generator, period]
expression: at(price, by=zone_of, over=zone, into=generator) * p <= 1000
What the declaration decides, the call may leave unsaid. Where the key has
one column and the key determines one column, the walk is the arrow the key
draws, and sum(p, by=gen_bus) and at(price, by=gen_bus) are complete:
sum consumes the key and produces the value, at consumes the value and
produces the key. Where a side has several candidates — two key columns, two
value columns — the call names it, and the refusal lists the candidates.
zone_of above has two key columns, so sum names over=, while into=zone
could have been left out.
A partition walks a key column and groups by the value columns.
shift(x, along=d, by=l), sum_back(x, along=d, by=l) and
position(d, by=l) take the one key column over d; the other key columns
are joined on, and the group is the value tuple. within= names the value columns the group is made of
where the table has several: shift(x, along=snapshot, by=cal, within=week)
walks within weeks of a calendar declared once over [snapshot, day, week],
and a value column not named is not read.
The rules, each decided at load with a refusal naming the rewrite:
over=andinto=name columns of the relationby=names, one each or a list each, and no column on both sides.into=is refused without aby=, since a column needs the table that holds it.over=without one names a dimension of the operand instead, which issum(p, over=period).sum(p, by=gen_bt, into=[bus, technology])lands one table with two value columns on the productbus × technologyin one join;sum(p, by=zone_of, over=[generator, period])consumes both key columns at once, which issum(sum(p, by=zone_of, over=generator), over=period)said once;at(tech_cap, by=gen_bt, over=[bus, technology])reads a two-column slot at each generator.- The operand carries every joined column's dimension, each once. The map is read at the key columns not walked, so there is no reading it at a coordinate that lacks them; two joined columns over one dimension have nothing to tell apart.
- A produced dimension the operand already carries is joined on too.
sum(load * p, by=gen_bus)withload[snapshot, bus]restricts each term to the row where the generator's bus is the row's bus — a masked sum, which is what the join says. atreads one value. Its key lies insideinto=and the joined columns, or the call is refused; a bare relation is never read byat.sumadds its rows up. So the reverse holds: asumwhose key lies insideinto=and the joined columns finds one row per coordinate and adds up nothing, which is a read — it is refused towardat.sumwalks to a value column;atwalks to the key.- A partition walks the one key column over the dimension it walks, and
groups by the value columns
within=names — all of them where it names none.within=naming a key column is refused, and a bare relation partitions nothing. The group may hold two columns over one dimension, a pair of buses say: a partition lands nothing, so nothing needs the dimension twice. - A
by=list walks each relation by its declared arrow.by=[a, b]is one grouping, so no column keyword has anything to name; every relation in it consumes the same dimension, joins on its own other columns, and no two produce the same dimension. - A
wherecomparison reads a value column of a keyed relation at its key.zone_of == 'north'reads the one value column;ends.bus0 != ends.bus1names the columns where there are several. The frame carries the key's dimensions, and two relations compared have keys over the same dimensions and columns over one. A bare name —where: gen_bus— tests that a row exists: at the key for a keyed relation, at every column for a bare relation. - Every column is over a declared dimension, every column name is distinct, the key names columns the relation has, and does not name all of them.
Every relation name joins the flat namespace, so a relation may not shadow a
dimension. generator's map onto bus is gen_bus, never a second bus.
Roles#
A list under columns: names each column after its dimension. Two columns over
one dimension need names of their own, and the mapping form gives them:
relations:
ends: { columns: { line: line, bus0: bus, bus1: bus }, key: line } # a line's two ends, one table
rep_of: { columns: { snapshot: snapshot, rep: snapshot }, key: snapshot } # the representative snapshot
sum(f, by=ends, over=line, into=bus1) - sum(f, by=ends, over=line, into=bus0)
is the nodal balance through one table where two relations did it before, and
where: "ends.bus0 != ends.bus1" excludes a self-loop by comparing two of its
columns.
rep_of relates a dimension to itself, which is how a clustered year is run
on a few typical days: every snapshot names the one that stands for it.
Nothing changes in the rules — snapshot is consumed and rep produced, both
over one dimension, so the frame is unchanged through sum(by=) and at(by=)
alike:
constraints:
representative: # every snapshot takes its representative's value
dims: [snapshot]
expression: p == at(p, by=rep_of)
weighted: # the snapshots a representative stands for, summed onto it
dims: [snapshot]
expression: sum(p, by=rep_of) <= 100
A self-map is directional exactly as far as its key says. key: snapshot
makes rep a function of snapshot, so the arrow runs from a snapshot to its
representative: at reads along it and sum collects against it, the inverse
of a many-to-one map being one-to-many, reachable as a grouping and never as
a function. Two steps along the arrow are two nested calls. Without a key the
same two columns are an undirected relation — a neighbour table — which sum
walks either way and nothing reads. Selecting the representatives themselves,
the rows where the map is the identity, is not a comparison the language has,
since a relation is never compared to a dimension; declare a bool parameter
for them.
How the map is supplied#
gen_bus is a source key like any other, carrying one column per column
declared, named after the column:
sources = {
'generator': ['g1', 'g2', 'g3'],
'gen_bus': pl.DataFrame({'generator': ['g1', 'g2'], 'bus': ['north', 'south']}),
}
A partial map is the rows it has. g3 is in no row, so g3 sits on no
bus — absence is the absent row, exactly as it is for a parameter, and a null
in any column is refused for saying both at once. A keyed table holds one row
per key tuple, and a value matching no label of its column's dimension is a
typo rather than a new member. Values are never inferred from the parameters
that use a dimension: inferring would let a mistyped label extend the label
set instead of being rejected.
Supplying it this way touches no table but its own, which is what a caller who did not generate the index needs: a model can be extended with a relation the same way it can be extended with a parameter. A column of a dimension's index named after a relation is refused rather than read — an index may carry any other extra, and this one would be a map read by accident.
Dimension, relation or parameter?#
Every column of data is one of the three. What decides which is what the math does with the column, not what the column holds:
| The column… | is declared as | because |
|---|---|---|
| is an axis: something is indexed by it, or an aggregation lands terms on it | a dimension |
its members are the coordinate set every table over it is reindexed onto |
| has one value per member of a dimension, or per tuple of several — a generator's bus, a line's two ends, a generator's zone by period | a relation with that key |
it is a map every operator walks, and its values are checked against the dimensions they name |
| relates members of two dimensions many-to-many, with nothing to weigh — which buses a generator may connect to | a relation with no key |
sum walks it with both ends named, and a bare where tests it. Nothing reads it, because there is no one value to read |
| relates members of two dimensions many-to-many, with a weight per pair — a link's efficiency to each bus, a cycle's lines | a parameter over both |
the weight is the data, its row set is the relation, and the aggregation is sum(w * x, over=a) |
| is a label set the model only selects on or counts within — a period, a season, a zone | a dimension, and a keyed relation onto it |
the membership check is worth one line and one member list |
| scales terms — a coefficient, a bound, an offset | a parameter (float or int) |
arithmetic is over numbers (dtype) |
| is a per-row attribute the math only selects on — a fuel, a constraint's sense | a str parameter |
it names rows rather than scaling them, and no set is declared to check its values against |
| is a mask | a bool parameter |
a bare name in a where is its own answer |
Two rules follow from the table. If b has one value per a, then b is a
relation keyed by a, and not a dimension: a dims product over two
dimensions that depend on each other, cut back with a mask, is the shape that
relations replaces.
And everything under dimensions: is an axis. A dimension is never legal where
a value belongs, because it is a coordinate space and not data. To use a
dimension's coordinates as data, declare a parameter over it.
python -m math_spec check advises on a declared dimension that nothing is
indexed by, nothing aggregates into and no relation has a column over
(errors).