GraphML vs GEXF in NetworkX: Choosing a Serialization Format for Weighted Graphs
A decision guide for picking GraphML or GEXF when serializing weighted NetworkX graphs: attribute fidelity, schema weight, downstream tooling, and a round-trip validation script.
30 Sept 2026, 22:05 UTC

You have a weighted graph in NetworkX and need to save it to disk so a colleague, another tool, or a future run of your pipeline can reload it with attributes intact. NetworkX ships built-in readers and writers for both GraphML and GEXF, and both preserve node and edge attributes. The practical difference is schema complexity, dependency footprint, and how each format represents your weights. This guide compares them and shows a round-trip validation you can run before committing to one.
The decision and its constraints
The decision: which XML-based exchange format should your NetworkX pipeline write? The constraints that usually matter:
- Attribute fidelity: weights and other node/edge attributes must survive a save/load cycle without type loss.
- Dependencies: whether you can rely only on the Python standard library.
- Downstream tools: whether the file will be opened in Gephi, another NetworkX script, or a generic XML consumer.
- Graph size: parse overhead grows with schema verbosity.
Supported options compared
| Property | GraphML | GEXF |
|---|---|---|
| NetworkX functions | nx.write_graphml / nx.read_graphml | nx.write_gexf / nx.read_gexf |
| Attribute encoding | <key> declarations plus <data> elements | <attvalue> elements under <attributes> |
| Extra metadata | Minimal; structure and data only | Richer: supports visualization-oriented metadata such as layout and styling concepts |
| Schema weight | Leaner schema, generally faster to parse | Larger schema, more verbose output |
| Typical consumer | General-purpose tools, yEd, many graph libraries | Gephi (GEXF is its native exchange format) |
| Type handling | Declared types via key attr.type | Declared types via attribute class and type attributes |
Trade-offs in practice
Both formats are XML and both are handled by NetworkX's standard-library-based parsers, so neither adds a hard dependency to a NetworkX install. The real trade-offs are elsewhere.
Choose GraphML when the file is a data artifact: you want the smallest, most portable representation of nodes, edges, and typed attributes, and the consumer is another script or a general graph tool. Its <key>/<data> model maps directly onto NetworkX's attribute dictionaries, and the leaner schema keeps parse time down on large graphs.
Choose GEXF when the file is headed for visualization, especially Gephi. GEXF was designed to carry presentation context alongside topology, so if a teammate will open the export and style it interactively, GEXF avoids a conversion step. The cost is a more verbose file and a larger schema surface where malformed input can break parsing.
Type fidelity caveat for both: XML is text, so values round-trip through declared types. Stick to int, float, str, and bool for attributes. Exotic Python objects (tuples, custom classes) are not reliably serializable in either writer; convert them to strings or numbers first. Also note that GraphML writers can raise errors on attribute names or types the schema cannot express — validate on a small sample before exporting a full graph.
Concrete implementation and round-trip validation
The safest way to decide is to verify that your actual graph survives both formats. Run this in any Python environment with NetworkX installed (no special permissions needed; it writes to a temp directory):
import tempfile, os
import networkx as nx
G = nx.DiGraph()
G.add_edge("a", "b", weight=1.5, label="ab")
G.add_edge("b", "c", weight=2.0, label="bc")
G.nodes["a"]["role"] = "source"
tmp = tempfile.mkdtemp()
for fmt, writer, reader in [
("graphml", nx.write_graphml, nx.read_graphml),
("gexf", nx.write_gexf, nx.read_gexf),
]:
path = os.path.join(tmp, f"g.{fmt}")
writer(G, path)
H = reader(G.__class__) if False else reader(path)
for u, v, data in G.edges(data=True):
for k, val in data.items():
assert H[u][v][k] == val, (fmt, u, v, k, H[u][v].get(k), val)
print(fmt, "round-trip OK")
What to check: the assertions compare every edge attribute between the original and reloaded graph. If a weight comes back as a string ("1.5" instead of 1.5), the assertion fails and you know that format/version combination loses the type for your data — in that case, coerce on load (float(H[u][v]["weight"])) or adjust how you write attributes. Extend the loop to node attributes if you rely on them.
Limitations and failure modes
- Namespace/version mismatches: GraphML files declaring a schema version the reader does not expect can fail to import. If you receive GraphML from another tool and
read_graphmlerrors, inspect the root element's namespace first. - Malformed GEXF: missing required attributes or non-conforming XML can make GEXF parsing fail; behavior varies across NetworkX versions, so pin your version in requirements and test after upgrades.
- Mixed-type attributes: if the same attribute key is an int on one edge and a string on another, both writers may refuse to infer a single type. Normalize types before export.
- Graph-level attributes and multigraphs: support for edge keys in multigraphs differs between formats; if you use
MultiDiGraph, include it in your round-trip test rather than assuming parity.
The practical rule: default to GraphML for data interchange, switch to GEXF when Gephi or visualization metadata is the destination, and in either case gate the choice on a round-trip assertion against your real attribute set.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.