Skip to main content

nimblend arrays

nimblend is the layer below nimopt. Its vocabulary is dimensions, labels, entries and alignment, and it contains no optimization term. A nimopt caller uses these names when reading a solution or inspecting what a model built.

Import from nimblend itself, never from a submodule. Never read an array's .index or .data or a domain's .codes. Those are the raw index matrix, the value buffer and the ravelled members of the layer below the array layer. Reading them bypasses the contract, and a test in this repository fails on any of them.

Do not assemble one either. SparseArray.from_canonical below takes an index matrix built by the caller, and that is the work of the array layer. No module of nimopt calls it, and a test enforces that. A Domain returns the array over its own members instead.

Array​

A labeled N-dimensional array. Array is the contract both implementations satisfy, and the interface a caller writes against.

absence declares the meaning of a coordinate the array does not have: "empty" that it contributes nothing, "unknown" that it was not modeled. Operators and reductions follow from that declaration. The declaration is part of the definition of the array, not a hint.

MemberReturns
dims, shape, nnzthe dimensions, their sizes, and the number of entries
coordsthe coordinate of each dimension, resolving its labels
absencethe meaning of a coordinate the array does not have
as_empty(), as_unknown()the array under the other absence declaration
values(), coordinates()the entries and their multi-indices
domain()the coordinates the array has
sum, min, max, meanreductions over named dimensions
sel, restricta selection by label, and by domain
rename, transpose, expand, conformreshaping the dimensions
shift, rolla lag that drops, and one that wraps
groupentries combined into a destination
to_dense(fill=None)the entries as an ndarray
+, -, *, /, unary -arithmetic over one frame, and with a scalar

The arithmetic is part of the contract, not an implementation's own: a caller writing coefficient * columns is writing against Array. Both implementations behave alike, including over frames that differ. One frame nested inside the other broadcasts over the wider; frames sharing some dimensions align on those and multiply out the rest. Frames sharing no dimension raise. A product mixing the two implementations returns a SparseArray: a product intersects presence and has at most the entries of the sparse operand.

Absence and zero remain distinct. A stored 0.0 is a coordinate present with the value zero. A coordinate the array does not have is a different case.

import numpy as np
import nimblend as nb

labels = {"A": np.array(["a0", "a1"]), "B": np.array(["b0", "b1"])}
array = nb.SparseArray.from_dense(np.array([[1.0, 0.0], [0.0, 2.0]]), labels)

print(isinstance(array, nb.Array))
print(array.dims, array.shape, array.absence)
print(array.nnz)
print(array.sum("B").to_dense())
Output
True
('A', 'B') (2, 2) empty
4
[1. 2.]

from_dense stores every cell it was given, and this array has four entries and not two. The zeros are stored, and a stored value is present.

SparseArray​

Entries in canonical order, under a coordinate per dimension. An entry that is not stored is absent.

It is what an operation returns when the result is sparse, and what Solution.primal returns for a variable over a subset: a dense frame there would be the grid the variable was declared to avoid.

MemberReturns
SparseArray.from_dense(values, labels)every cell of an ndarray
SparseArray.from_canonical(index, values, coords, dims)entries already in order
as_empty(), as_unknown()the same entries under the other declaration
to_csr()the entries as compressed rows
import numpy as np
import nimblend as nb

labels = {"A": np.array(["a0", "a1"]), "B": np.array(["b0", "b1"])}
array = nb.SparseArray.from_dense(np.array([[1.0, 0.0], [0.0, 2.0]]), labels)

print(array.absence)
print(array.as_unknown().absence)
print(array.to_dense())
Output
empty
unknown
[[1. 0.]
[0. 2.]]

DenseArray​

An ndarray over labeled dimensions, distinguishing absence from zero.

The storage of presence follows the absence declaration: the two declarations require opposite behavior from an operator. An "unknown" array marks absence with NaN. NaN propagates through arithmetic at no cost and requires no storage beside the values. An "empty" array stores a boolean mask: absence is the additive identity there, and substituting it costs less than marking it.

Solution.primal returns one of these for a variable over a full product. The solver returns a value at every cell of the frame in column order, and the values reshape with no index built.

import numpy as np
import nimblend as nb

coords = {
"A": nb.StoredCoord(np.array(["a0", "a1"])),
"B": nb.StoredCoord(np.array(["b0", "b1"])),
}
array = nb.DenseArray(np.array([[1.0, 0.0], [0.0, 2.0]]), coords, ("A", "B"))

print(array.absence, array.nnz)
print(array.present)
print(array.as_unknown().absence)
Output
empty 4
[[ True True]
[ True True]]
unknown

Building an array from columns​

nimblend itself provides the two constructors a caller uses when the data is not already an ndarray.

ConstructorReturns
from_long(dims, coords, labels, values)one label column per dimension and one value column
from_dense(values, labels)every cell of an ndarray
is_canonical(index, shape)whether buffers are in the order from_canonical takes

from_long resolves each label through the coordinate that dimension already has. A caller with coordinates of its own therefore gives its entries as labels and resolves no position itself. The columns are read in parallel and must have equal length.

import numpy as np
import nimblend as nb

coords = {
"t": nb.StoredCoord(np.array([2030, 2040])),
"r": nb.StoredCoord(np.array(["DE", "FR"])),
}
arr = nb.from_long(
("t", "r"),
coords,
{"t": np.array([2030, 2040, 2040]), "r": np.array(["DE", "DE", "FR"])},
np.array([5.0, 6.0, 7.0]),
)
print(arr.nnz, arr.to_dense()[1, 1])
Output
3 7.0

A column of a different length raises ValueError; the message gives the column and both lengths.

import numpy as np
import nimblend as nb

nb.from_long(
("t",),
{"t": nb.StoredCoord(np.array([2030, 2040]))},
{"t": np.array([2030, 2040])},
np.array([1.0]),
)
Raises ValueError
ValueError: label columns have lengths {'t': 2} and the value column has length 1; pass columns of equal length

is_canonical reports whether buffers meet the precondition of SparseArray.from_canonical: entries sorted by ravel key, with no repeat. Verifying it inside from_canonical would cost the ravel that this path avoids.

import numpy as np
import nimblend as nb

ordered = np.array([[0, 0, 1], [0, 1, 0]], dtype=np.int32)
print(nb.is_canonical(ordered, (2, 2)))
print(nb.is_canonical(ordered[:, ::-1].copy(), (2, 2)))
Output
True
False

The frame of a binary result​

combined_dims(left, right) returns the dimensions a binary operator's result has, from the dimensions of the two operands alone. A caller reads it before materialising either operand. A combination therefore reports its frame while its data is unbound.

OperandsResult
equal framesthat frame, in its order
one frame nested in the otherthe wider
frames that overlapthe left, then the dimensions only the right has
frames sharing no dimensionraises
import nimblend as nb

print(nb.combined_dims(("P", "Q"), ("Q", "R")))
print(nb.combined_dims(("P",), ("P", "Q")))
Output
('P', 'Q', 'R')
('P', 'Q')

Frames sharing no dimension have no common dimension to align on, and combined_dims raises. Their combination would be an outer product.

import nimblend as nb

nb.combined_dims(("P",), ("Q",))
Raises ValueError
ValueError: frames ('P',) and ('Q',) share no dimension; pass operands that share a dimension

Adding several arrays​

sum_arrays(arrays) adds several arrays over one frame in one merge. The result equals adding them in order with +. A chain of + merges the running sum once per operand, and sum_arrays sorts the entries of all operands once. Under "empty" the result has every coordinate of any operand. Under "unknown" it has only the coordinates of every operand. The operands have the same dimensions, labels and absence.

import numpy as np
import nimblend as nb

labels = {"A": np.array(["a0", "a1"]), "B": np.array(["b0", "b1"])}
first = nb.SparseArray.from_dense(np.array([[1.0, 0.0], [0.0, 2.0]]), labels)
second = nb.SparseArray.from_dense(np.array([[10.0, 20.0], [30.0, 40.0]]), labels)

print(nb.sum_arrays([first, second, first]).to_dense())
Output
[[12. 20.]
[30. 44.]]

Densifying an unknown array​

An array declaring "unknown" that does not have every coordinate of its frame raises on to_dense() without a fill. There is no value it can place at the rest, and choosing one silently would invent a value.

import numpy as np
import nimblend as nb

labels = {"A": np.array(["a0", "a1"]), "B": np.array(["b0", "b1"])}
partial = nb.SparseArray.from_canonical(
np.array([[0], [0]], dtype=np.int32),
np.array([1.0]),
{
"A": nb.StoredCoord(labels["A"]),
"B": nb.StoredCoord(labels["B"]),
},
("A", "B"),
absence="unknown",
)

partial.to_dense()
Raises ValueError
ValueError: absence is 'unknown' and the array has no value at 3 of 4 coordinates; pass fill=<value> to to_dense()

With a fill value, the grid is returned.

import numpy as np
import nimblend as nb

labels = {"A": np.array(["a0", "a1"]), "B": np.array(["b0", "b1"])}
partial = nb.SparseArray.from_canonical(
np.array([[0], [0]], dtype=np.int32),
np.array([1.0]),
{
"A": nb.StoredCoord(labels["A"]),
"B": nb.StoredCoord(labels["B"]),
},
("A", "B"),
absence="unknown",
)

print(partial.to_dense(fill=np.nan))
Output
[[ 1. nan]
[nan nan]]