Skip to content

Block and Frame

The chemistry is finished. You need distances, an energy, or a file the engine will accept. How do you leave the molecular graph without losing structure?

Engines do not read graphs. They read aligned tables: coordinates, types, and index-based topology. MolPy makes that hand-off explicit.

A Block is one columnar table (every column is a NumPy array over the same rows). A Frame is a named set of blocks plus metadata — one complete snapshot of the system.

What they are not: an editable bond graph. Editing chemistry stays on Atomistic. Convert with to_frame() when the structure is stable.

Why two representations?

Build and edit needs graph traversal (SMILES, polymer growth, reactions). Compute and export needs arrays (distances, RDF, LAMMPS data). Hiding the conversion invites silent mismatch; keeping both representations keeps the boundary honest.

A block answers “what are the atoms?” or “what are the bonds?” — one table per kind of data. A frame answers “what is the full state right now?” by grouping those tables (and a box, when present).

Block: a columnar table backed by NumPy

Pass a dictionary of array-like values; each value becomes a NumPy array.

import molpy as mp
import numpy as np

atoms = mp.Block({
 "element": ["O", "H", "H"],
 "x": [0.000, 0.957, -0.239],
 "y": [0.000, 0.000, 0.927],
 "z": [0.000, 0.000, 0.000],
})

print(atoms.n_rows) # 3
print(list(atoms.keys())) # ['element', 'x', 'y', 'z']

Reading a column returns an np.ndarray. This means all of NumPy is immediately available — no conversion step, no special accessor.

print(atoms["x"].dtype) # float64
print(atoms["element"].dtype) # <U1 (Unicode string)

A common pattern is stacking numeric columns into a 2D array for vectorized computation.

xyz = atoms[["x", "y", "z"]] # shape (3, 3)
r = np.linalg.norm(xyz, axis=1)
print(r)

Row selection returns a new Block

Slicing, boolean masks, and fancy indexing all produce a new Block. The original is never modified.

hydrogens = atoms[atoms["element"] == "H"]
print(hydrogens.n_rows) # 2
print(hydrogens["x"]) # [ 0.957 -0.239]

first_two = atoms[0:2]
print(first_two["element"]) # ['O' 'H']

If you need a single scalar value, index the column first, then the row.

print(atoms["x"][0]) # 0.0

Adding and removing columns

Setting a key inserts or overwrites a column. Deleting a key removes it. Both operations follow standard Python mapping conventions.

atoms_with_r = atoms.copy()
atoms_with_r["r"] = np.linalg.norm(atoms_with_r[["x", "y", "z"]], axis=1)
print(list(atoms_with_r.keys())) # ['element', 'x', 'y', 'z', 'r']

del atoms_with_r["r"]
print(list(atoms_with_r.keys())) # ['element', 'x', 'y', 'z']

Renaming columns

Block.rename() changes a column key in place and keeps the column where it was. FieldFormatter uses it to translate between format-specific and canonical field names at an I/O boundary.

b = mp.Block({"q": [0.1, -0.2], "x": [1.0, 2.0]})
b.rename("q", "charge")
print(list(b.keys())) # ['charge', 'x'] — the renamed column keeps its place

Copy semantics

Block.copy() is a deep copy: every column gets its own buffer. Writing into the copy never reaches the original.

independent = atoms.copy()
independent["x"][0] = 999.0
print(atoms["x"][0]) # 0.0 — original unchanged

Avoiding mutation

The idiomatic MolPy pattern is to avoid in-place array mutation entirely. Instead of modifying a column, assign a new array: block["x"] = block["x"] + 1.0. This always produces an independent column and is consistent with MolPy's immutable-data philosophy.

Frame: a named collection of Blocks

A molecular system usually needs more than one table. Atom coordinates are one table, bond indices are another, and the snapshot itself has metadata — a timestep, a description, provenance. Frame groups all of that into one object.

frame = mp.Frame({
 "atoms": mp.Block({
 "element": ["O", "H", "H"],
 "x": [0.000, 0.957, -0.239],
 "y": [0.000, 0.000, 0.927],
 "z": [0.000, 0.000, 0.000],
 }),
 "bonds": mp.Block({
 "atomi": [0, 0],
 "atomj": [1, 2],
 }),
})
frame.meta = {
 "timestep": 0,
 "description": "water",
}

frame.meta is a live mapping in insertion order: get a Python scalar, set a Python scalar. Exact dtypes stay in the store — frame.meta.dtype(key) reports one, frame.meta.typed() hands out every value as an mp.core.MetaValue — and you pass mp.core.MetaValue only when you need a specific one.

print(frame.meta["timestep"]) # 0
print(frame.meta["description"]) # water
print(frame.meta.dtype("timestep")) # i64

What comes back is frozen: a vector or a JSON array is a tuple, and a JSON object is a read-only MetaDocument. To change a nested value, copy, edit and store it back:

frame.meta["run"] = {"step": 0, "ensemble": "nvt"}
run = frame.meta["run"].copy() # a plain dict
run["step"] = 3
frame.meta["run"] = run
print(frame.meta["run"]["step"]) # 3

Accessing a block by name returns a Block — a handle on the stored table, so every column operation works the same way and a write lands in the frame.

atoms = frame["atoms"]
print(atoms["x"]) # [ 0.     0.957 -0.239]

You can add, replace, or delete blocks at any time.

frame["tags"] = {"label": ["oxygen", "hydrogen", "hydrogen"]}
print(type(frame["tags"]) is mp.Block) # True

del frame["tags"]
print("tags" in frame) # False

Box is a first-class attribute

A periodic simulation cell is attached directly to frame.box, not stored in metadata. This ensures Frame.copy() preserves the box and I/O round-trips work correctly.

frame.box = mp.Box.cube(20.0)
print(frame.box.lengths) # [20. 20. 20.]

# copy() preserves box
frame2 = frame.copy()
print(frame2.box.lengths) # [20. 20. 20.]

frame.box is None when no box has been assigned (e.g., for isolated molecules).

Round-trips through plain dictionaries

dict(block) is its columns, and mp.Block(mapping) builds a table back. A frame's parts are exactly its constructor arguments: the blocks, meta, and box. (Frames and blocks also pickle, masks and metadata dtypes included.)

columns = {name: dict(frame[name]) for name in frame.keys()}
print(sorted(columns)) # ['atoms', 'bonds']

restored = mp.Frame(columns, meta=dict(frame.meta), box=frame.box)
print(sorted(restored.keys())) # ['atoms', 'bonds']

When Block and Frame are the right choice

Use Atomistic when you still need to edit the molecular graph — adding atoms, defining bonds, querying connectivity. Use Block and Frame when the chemistry is settled and your next task involves arrays, export, or analysis.

The two representations can coexist. Many workflows keep an Atomistic around for reference while producing frames for numerical work. The important thing is knowing which object carries the meaning you care about at each stage.

Once your system lives in a periodic cell, coordinates alone are not enough — distances depend on the simulation box. That is the subject of the next page.

See also: Atomistic and Topology, Box and Periodicity.