Skip to content

Notation

Chemical string notation, parsed by the native core. Three notations are supported — SMILES, CGsmiles and SMARTS — and each is reached through a type; the reader functions are mp.io.read_smiles_str and mp.io.read_cgsmiles_str, a whole molecule straight to a graph. Every notation type is a native type in the molpy module that mirrors its molrs owner — SMILES in mp.io.smiles (mp.io.smiles.SmilesIr), CGsmiles in mp.io.cgsmiles (mp.io.cgsmiles.CgSmilesIr), SMARTS in mp.perceive (mp.perceive.SmartsPattern).

Quick reference

Expression Input Output Use when
mp.io.read_smiles_str(s) SMILES Atomistic One molecule (a .-separated set is refused)
mp.io.smiles.SmilesIr(s).to_atomistic() SMILES Atomistic Any SMILES (several components: one disconnected graph)
mp.io.smiles.SmilesIr(s) SMILES SmilesIr Inspect before converting
mp.io.smiles.SmilesIr(s).n_components SMILES int How many molecules the string names
mp.io.smiles.SmilesIr(s).components() dot-separated SMILES list[Atomistic] One graph per component ([Li+].[F-])
mp.perceive.SmartsPattern(p) SMARTS SmartsPattern Pattern matching / typification

Canonical example

import molpy as mp

mol = mp.io.read_smiles_str("CCO") # Atomistic (heavy atoms only)
mol = mp.perceive.add_hydrogens(mol) #... with hydrogens

ions = mp.io.smiles.SmilesIr("[Li+].[F-]").components() # [Atomistic, Atomistic]

query = mp.perceive.SmartsPattern("[C;X4][O;H1]") # compiled query
query.find_matches(mol) # -> list[SmartsMatch]

A .-separated string names a set of molecules: mp.io.read_smiles_str refuses it, SmilesIr(s).to_atomistic() returns them as one disconnected graph, components() one graph each, and n_components says how many there are.

Polymer notations

CGsmiles is parsed by mp.io.cgsmiles.CgSmilesIr: templates() gives each fragment as an Atomistic whose bonding descriptors are ports (one fragment body alone: mp.io.smiles.SmilesIr.from_fragment(body).to_template()), and to_coarsegrain() gives the site graph that mp.builder.Assembler builds. BigSMILES and G-BigSMILES are not parsed.

  • mp.perceive.add_hydrogens, assign_aromaticity, assign_rings, assign_stereo — hydrogens, aromaticity, rings, stereo (perceive before you match: X4 and H1 count what is actually in the graph)
  • mp.perceive.RingSet — ring / ring-system queries
  • mp.perceive.Reaction — a reaction SMARTS applied to a graph in place: forms and breaks bonds, deletes the unmapped leaving atoms
  • Guide: Parsing Chemistry

Full API

read_smiles_str

read_smiles_str(smiles)

One molecule from a SMILES string: connectivity only, no implicit H, no coordinates. A '.'-separated set raises SmilesError (a ValueError) naming SmilesIr(s).components().

SmilesIr

SmilesIr(smiles)

Intermediate representation of a parsed SMILES string (or SMILES fragment body).

to_atomistic() is the plain conversion: it refuses SMARTS query atoms and, since it will not drop them silently, any node carrying a bonding descriptor — which is what a CgFragmentDef.body from the last CGsmiles block holds. Build such a body's ported unit with to_template() (parse it with SmilesIr.from_fragment), or expand a whole string through CgSmilesIr.to_atomistic.

from_fragment classmethod

from_fragment(body)

Parse a CGsmiles fragment body (SMILES plus bonding descriptors, e.g. "[<]OCC[>]"); the plain constructor refuses descriptors.

to_template

to_template()

The ported unit of this body: heavy atoms plus one hydrogen handle and one port per bonding descriptor; no coordinates, no frag_id. SmilesIr.from_fragment("[<]OCC[>]").to_template() equals CgSmilesIr("{[#EO]}.{#EO=[<]OCC[>]}").templates()["EO"].

CgSmilesIr

CgSmilesIr(text)

Intermediate representation of a parsed CGsmiles string.

Constructing it parses, validates, expands and resolves the whole string; the value is then read, not built. levels are the resolution levels (coarsest first), fragments[k] resolves the names of levels[k], and to_atomistic() expands the lowest level into a topology-only graph whose atoms carry frag_id.

templates

templates()

One ported :class:Atomistic template per definition of the last fragment table, keyed by name. One body alone: SmilesIr.from_fragment(body).to_template().

to_coarsegrain

to_coarsegrain()

The coarsest level, levels[0], as a bead graph: one bead per node (bead_type only, no coordinates or mass), one CG bond per edge.

Raises

SmilesError (a ValueError) the IR breaks a reader invariant; no parsed string reaches this.

SmartsPattern

SmartsPattern(smarts)

Compiled, atom-map-aware SMARTS query over an :class:Atomistic.

Wraps the core Rust SMARTS engine (non-uniquified, RDKit uniquify=False). Daylight atom maps ([C:1]) add no match constraint; a match's :attr:SmartsMatch.mapping is its {map_number: atom_handle} dict.

from_environment classmethod

from_environment(
    mol,
    center,
    *,
    reach=1,
    atomic_number=True,
    include_degree=True,
    include_h_count=True,
    include_charge=True,
    include_aromatic=True,
    include_ring_membership=False,
    include_ring_size=False,
    include_explicit_h_atoms=False,
    include_bond_orders=True,
    neighbor_style="chain",
    canonical_neighbor_order=True,
)

The pattern that states the local environment of center in mol out to reach bonds; it matches mol at center.

SmartsMatch

One SMARTS embedding.

SmilesError

Bases: ValueError

Raised when a SMILES / SMARTS / CGsmiles string is refused.

Subclasses ValueError; str(e) is the message the Rust error renders, caret line included. The four attributes are set on every instance.

Attributes

kind : str Variant name of the rule that was broken, payload dropped — "UnclosedBranch", "UnexpectedEnd", "CgNotExpandable", ... span : tuple[int, int] Byte range of the offending text within input; the end is clamped to len(input), since the scanner reports end-of-input one byte past the text. input : str The offending string, empty for errors raised past the parser (the expansion and emit stages are handed an IR, not the text). notation : str Which notation was being read or written: "smiles", "smarts" or "cgsmiles".