Block and Frame¶
The chemistry is finished. You need distances, an energy, or a file the engine will accept. How do you leave the molecular graph without losing structure?
Engines do not read graphs. They read aligned tables: coordinates, types, and index-based topology. MolPy makes that hand-off explicit.
A Block is one columnar table (every column is a NumPy array over the same
rows). A Frame is a named set of blocks plus metadata — one complete
snapshot of the system.
What they are not: an editable bond graph. Editing chemistry stays on
Atomistic. Convert with to_frame() when the
structure is stable.
Why two representations?¶
Build and edit needs graph traversal (SMILES, polymer growth, reactions). Compute and export needs arrays (distances, RDF, LAMMPS data). Hiding the conversion invites silent mismatch; keeping both representations keeps the boundary honest.
A block answers “what are the atoms?” or “what are the bonds?” — one table per kind of data. A frame answers “what is the full state right now?” by grouping those tables (and a box, when present).
Block: a columnar table backed by NumPy¶
Pass a dictionary of array-like values; each value becomes a NumPy array.
import molpy as mp
import numpy as np
atoms = mp.Block({
"element": ["O", "H", "H"],
"x": [0.000, 0.957, -0.239],
"y": [0.000, 0.000, 0.927],
"z": [0.000, 0.000, 0.000],
})
print(atoms.n_rows) # 3
print(list(atoms.keys())) # ['element', 'x', 'y', 'z']
Reading a column returns an np.ndarray. This means all of NumPy is immediately available — no conversion step, no special accessor.
A common pattern is stacking numeric columns into a 2D array for vectorized computation.
Row selection returns a new Block¶
Slicing, boolean masks, and fancy indexing all produce a new Block. The original is never modified.
hydrogens = atoms[atoms["element"] == "H"]
print(hydrogens.n_rows) # 2
print(hydrogens["x"]) # [ 0.957 -0.239]
first_two = atoms[0:2]
print(first_two["element"]) # ['O' 'H']
If you need a single scalar value, index the column first, then the row.
Adding and removing columns¶
Setting a key inserts or overwrites a column. Deleting a key removes it. Both operations follow standard Python mapping conventions.
atoms_with_r = atoms.copy()
atoms_with_r["r"] = np.linalg.norm(atoms_with_r[["x", "y", "z"]], axis=1)
print(list(atoms_with_r.keys())) # ['element', 'x', 'y', 'z', 'r']
del atoms_with_r["r"]
print(list(atoms_with_r.keys())) # ['element', 'x', 'y', 'z']
Renaming columns¶
Block.rename() changes a column key in place and keeps the column where it was. FieldFormatter uses it to translate between format-specific and canonical field names at an I/O boundary.
b = mp.Block({"q": [0.1, -0.2], "x": [1.0, 2.0]})
b.rename("q", "charge")
print(list(b.keys())) # ['charge', 'x'] — the renamed column keeps its place
Copy semantics¶
Block.copy() is a deep copy: every column gets its own buffer. Writing into
the copy never reaches the original.
independent = atoms.copy()
independent["x"][0] = 999.0
print(atoms["x"][0]) # 0.0 — original unchanged
Avoiding mutation
The idiomatic MolPy pattern is to avoid in-place array mutation entirely. Instead of modifying a column, assign a new array: block["x"] = block["x"] + 1.0. This always produces an independent column and is consistent with MolPy's immutable-data philosophy.
Frame: a named collection of Blocks¶
A molecular system usually needs more than one table. Atom coordinates are one table, bond indices are another, and the snapshot itself has metadata — a timestep, a description, provenance. Frame groups all of that into one object.
frame = mp.Frame({
"atoms": mp.Block({
"element": ["O", "H", "H"],
"x": [0.000, 0.957, -0.239],
"y": [0.000, 0.000, 0.927],
"z": [0.000, 0.000, 0.000],
}),
"bonds": mp.Block({
"atomi": [0, 0],
"atomj": [1, 2],
}),
})
frame.meta = {
"timestep": 0,
"description": "water",
}
frame.meta is a live mapping in insertion order: get a Python scalar, set a Python scalar. Exact dtypes stay in the store — frame.meta.dtype(key) reports one, frame.meta.typed() hands out every value as an mp.core.MetaValue — and you pass mp.core.MetaValue only when you need a specific one.
print(frame.meta["timestep"]) # 0
print(frame.meta["description"]) # water
print(frame.meta.dtype("timestep")) # i64
What comes back is frozen: a vector or a JSON array is a tuple, and a JSON object is a read-only MetaDocument. To change a nested value, copy, edit and store it back:
frame.meta["run"] = {"step": 0, "ensemble": "nvt"}
run = frame.meta["run"].copy() # a plain dict
run["step"] = 3
frame.meta["run"] = run
print(frame.meta["run"]["step"]) # 3
Accessing a block by name returns a Block — a handle on the stored table, so
every column operation works the same way and a write lands in the frame.
You can add, replace, or delete blocks at any time.
frame["tags"] = {"label": ["oxygen", "hydrogen", "hydrogen"]}
print(type(frame["tags"]) is mp.Block) # True
del frame["tags"]
print("tags" in frame) # False
Box is a first-class attribute¶
A periodic simulation cell is attached directly to frame.box, not stored in metadata. This ensures Frame.copy() preserves the box and I/O round-trips work correctly.
frame.box = mp.Box.cube(20.0)
print(frame.box.lengths) # [20. 20. 20.]
# copy() preserves box
frame2 = frame.copy()
print(frame2.box.lengths) # [20. 20. 20.]
frame.box is None when no box has been assigned (e.g., for isolated molecules).
Round-trips through plain dictionaries¶
dict(block) is its columns, and mp.Block(mapping) builds a table back. A
frame's parts are exactly its constructor arguments: the blocks, meta, and
box. (Frames and blocks also pickle, masks and metadata dtypes included.)
columns = {name: dict(frame[name]) for name in frame.keys()}
print(sorted(columns)) # ['atoms', 'bonds']
restored = mp.Frame(columns, meta=dict(frame.meta), box=frame.box)
print(sorted(restored.keys())) # ['atoms', 'bonds']
When Block and Frame are the right choice¶
Use Atomistic when you still need to edit the molecular graph — adding atoms, defining bonds, querying connectivity. Use Block and Frame when the chemistry is settled and your next task involves arrays, export, or analysis.
The two representations can coexist. Many workflows keep an Atomistic around for reference while producing frames for numerical work. The important thing is knowing which object carries the meaning you care about at each stage.
Once your system lives in a periodic cell, coordinates alone are not enough — distances depend on the simulation box. That is the subject of the next page.
See also: Atomistic and Topology, Box and Periodicity.