Abstract
Electron microscopy (EM) reveals atomic-scale structures that underpin catalysis, energy storage, and semiconductor reliability, yet current workflows remain fragmented across segmentation, crystallographic reconstruction, property modeling, and literature review, often requiring weeks of expert effort. Although recent artificial intelligence models have assisted individual steps, the diversity of EM modalities and tasks means existing approaches remain siloed and perform poorly in complex multistage workflows. We present EMSeek, a modular, provenance-tracked multiagent platform that connects EM to materials insight through five key units: reference-guided one-for-all segmentation, mask-aware reconstruction of crystal structures from EM data, a gated mixture of experts property predictor with uncertainty calibration, literature retrieval with citation anchoring, and physical consistency checks with audit-ready reporting. These units are orchestrated by large language models (LLMs) that automatically plan, invoke, and execute tools, minimizing human intervention. On 20 material systems and five tasks, EMSeek delivers segmentation about twice as fast as Segment Anything with higher accuracy, achieves more than 90% structural similarity on STEM2Mat, and, with about 2% labeled calibration, matches or surpasses strong single experts on three out-of-distribution property benchmarks. A complete query runs in 2 to 5 minutes per image, roughly 50 times faster than expert workflows. Case studies on two-dimensional lattices and nanoparticles validate EMSeek’s ability to automate complex workflows, with integrated uncertainty calibration and audit signals that provide scientists with rigorous yet actionable guidance to accelerate materials discovery.
EMSeek turns electron micrographs into masks, CIFs, and property predictions in minutes for faster materials discovery.
INTRODUCTION
Electron microscopy (EM) (1, 2) provides an unparalleled window into the atomic world, enabling direct visualization of defects (3, 4), lattice distortions (5–7), and chemical heterogeneity (8, 9) that dictate performance in catalysts, batteries, and semiconductors (10). The rapid proliferation of high-throughput imaging, variable-voltage instruments, and in situ experiments now produces immense volumes of data. However, most of these datasets remain underanalyzed: Advanced measurements lie dormant not because of their lack of significance but because expert interpretation is slow, fragmented, and irreproducible (11). Closing this gap calls for a previously unexplored paradigm that expands expertise, streamlines workflows, and unlocks the full scientific value of EM.
Recent advances in artificial intelligence (AI) (12–17) suggest such a transformation. Agentic systems have already accelerated software engineering, legal reasoning, and drug discovery by automating repetitive tasks (17–23), boosting productivity and enabling breakthroughs once deemed out of reach. However, in the context of EM, the diversity of imaging modalities and analysis tasks has limited the impact of AI to isolated steps. Existing approaches often remain siloed, assisting with lattice segmentation, structure indexing, or defect classification, but struggling to generalize across complex, multistage workflows. For example, AtomAI (24) contributes pixel-level methods for atomic segmentation, while AutoMat (25) advances automated approaches for structure indexing, each addressing distinct aspects of the EM analysis pipeline. This fragmentation means that raw micrographs are rarely connected to crystallographic models, property predictions, or literature evidence, leaving the “last mile” from observation to insight unresolved. This progress therefore prompts a pivotal question: Can we build a virtual EM scientist that autonomously tackles diverse imaging tasks, spans multiple materials subfields, and integrates insights across disciplines? An agent capable of managing thousands of concurrent micrographs could markedly elevate human productivity and hasten materials innovation.
Early attempts have relied on narrowly tailored pipelines: Deep-learning models for lattice segmentation (26) separate scripts for defect statistics (24) and ad hoc utilities for structure indexing (24). These workflows lack task-agnostic generalizability, struggling with low-contrast images, unfamiliar chemistries, and drifting beam conditions. They do not provide an integrated bridge from perception to property and knowledge. Building a truly general agent for EM therefore requires tight coupling of advanced reasoning with specialized actions, invoked dynamically and executed flexibly rather than through rigid scripts.
We present EMSeek, a modular and provenance-tracked multiagent platform that integrates perception, structure reconstruction, property inference, and literature reasoning into a unified workflow for EM. Instead of relying on a monolithic deep-learning stack, EMSeek delegates specialized tasks to a hierarchy of agents, coordinated by large language models (LLMs) that automatically plan, invoke, and execute tools with minimal human intervention. The platform comprises five key units. SegMentor performs reference-guided one-for-all segmentation, producing atom- and particle-level masks across diverse materials and imaging conditions. CrystalForge (EM2CIF) constrains reciprocal-space search to the masked regions and combines database retrieval with generative candidates to reconstruct density functional theory (DFT)–ready unit cells, enabling structure inference even for previously unseen chemistries. MatProphet serves as a gated mixture-of-experts (MoE) ensemble that fuses outputs from multiple interatomic models, calibrated with only 2% labeled data, to deliver properties such as formation energies and defect energetics together with uncertainty estimates. ScholarSeeker augments this workflow by retrieving and synthesizing evidence from thousands of domain articles, returning citation-anchored answers that reduce hallucinations. Last, Guardian validates physical plausibility, unit consistency, and provenance at each handoff, while Scribe compiles masks, crystallographic information files, property tables, and citations into an auditable report.
Benchmarks and case studies demonstrate that EMSeek generalizes across tasks and materials while substantially reducing turnaround time. On 20 material systems and five canonical tasks, SegMentor achieves segmentation at about twice the throughput of Segment Anything with higher accuracy, especially in low-contrast and drifted scenes. CrystalForge (EM2CIF) delivers more than 90% structural similarity on the STEM2Mat benchmark, with hybrid retrieval-plus-generation strategies lowering root mean square deviation (RMSD) and lattice error relative to pixel-level baselines. MatProphet, calibrated with ~2% labeled structures, matches or surpasses strong single experts on three out-of-distribution suites, including Matbench-Discovery, JARVIS-C2DB, and LeMat-Bulk, while providing calibrated uncertainty intervals. ScholarSeeker improves correctness and completeness on materials question-answering tasks, lowering hallucination rates compared with LLM and LLM + WebSearch baselines. A complete query, from raw micrograph to masks, CIF, properties, and cited answers, runs in 2 to 5 min per image, roughly 50 times faster than expert workflows. Case studies on two-dimensional (2D) lattices and nanoparticle size analysis demonstrate EMSeek’s capability to automate complex workflows. EMSeek thus represents a scalable framework for EM agents, positioning virtual scientists alongside human researchers to accelerate discovery from fundamental characterization to device optimization. An overview of EMSeek’s end-to-end multiagent workflow and the responsibilities of each agent is shown in Fig. 1. Our main contributions are as follows:
Fig. 1. Interactive EMSeek multiagent framework for end-to-end EM analysis.
(A) Conventional practice, from material synthesis and transmission electron microscopy (TEM) acquisition to data interpretation, demands thousands of manual actions (clicking Bragg disks, masking atoms, mapping strain regions of interest, picking photoluminescence (PL) peaks, plotting figures, and reading papers) and often takes weeks. (B) EMSeek restructures the pipeline into seven steps: (1) atomic/defect segmentation, (2) lattice reconstruction, (3) zero-shot property prediction, (4) literature retrieval, (5) physical-consistency checks, (6) statistics and visualization, and (7) report assembly and publishing, reducing turnaround to minutes. (C) EMSeek. After an EM image and objective are provided, the Maestro Agent assigns subtasks to SegMentor (segmentation), CrystalForge (structure reconstruction), MatProphet (property prediction with MoE-orchestrated models), ScholarSeeker (literature-augmented querying), and AnalyzerHub (quantitative analysis and plotting). The Guardian Agent verifies consistency, units, and provenance, and the Scribe Agent compiles figures, tables, and citations into a concise report.
1) EMSeek provides a modular, provenance-tracked multiagent system that unifies perception, structure modeling, property inference, and literature reasoning into a reproducible workflow for EM.
2) SegMentor introduces reference-guided one-for-all segmentation that produces atom- and particle-level masks across 20 material systems and five tasks, operating at roughly twice the throughput of Segment Anything with higher accuracy under noise and drift.
3) CrystalForge (EM2CIF) performs mask-aware crystallographic reconstruction by combining database retrieval with generative candidates, achieving more than 90% S.S. on STEM2Mat.
4) MatProphet uses a gated mixture-of-experts ensemble calibrated with only 2% labeled data, matching or surpassing strong single experts on out-of-distribution property prediction benchmarks while providing calibrated uncertainty estimates.
5) ScholarSeeker grounds results in peer-reviewed literature by retrieving and synthesizing evidence, yielding citation-anchored answers that improve correctness and completeness while reducing hallucinations.
6) Case studies on 2D lattices and on supported nanoparticles validate the generalization of EMSeek and highlight its potential to accelerate materials discovery and assist both scientists and nonexperts.
RESULTS
EMSeek accelerates EM-driven materials analysis
EM has long suffered an acute “perception-to-property” gap: Converting nanometer-scale images into quantitative structure and functional insight typically demands ad hoc segmentation, crystallographic indexing, and bespoke modeling, which can consume weeks of expert effort (27, 28). Inspired by recent agentic architectures (17–23), we recast this workflow as a controller-driven decision process with explicit quality checks and optional reruns. EMSeek delegates every stage, from raw micrograph to final property report to a constellation of specialized AI-based agents whose actions are expressed uniformly as executable code. A high-level Maestro Agent interprets the user’s question, decomposes it into subtasks, and supervises their execution, while a Guardian Agent safeguards provenance and data integrity. By replacing serial manual intervention with autonomous, self-documenting microservices, EMSeek compresses the entire analysis timeline to minutes and renders advanced EM analytics accessible through a single upload-and-query portal. EMSeek does not execute a fixed, linear pipeline. Instead, it follows a constrained interaction graph, where downstream agents can only consume specific artifacts produced upstream (for example, segmentation masks for structure reconstruction and CIFs for property prediction). This design limits workflow-level uncertainty propagation to a small set of permitted links and prevents unsupported shortcuts across stages. A detailed path enumeration is provided in the Supplementary Materials.
To achieve this acceleration, we built four core modules that tackle long-standing bottlenecks while explicitly balancing speed versus accuracy and flexibility versus computational cost. SegMentor performs reference-guided, one-for-all atom segmentation; its lightweight Ref-UNet (259 GFLOPs) surpasses the Segment Anything Model (560 GFLOPs) yet runs roughly twice as fast on a single graphics processing unit (GPU). CrystalForge (EM2CIF) turns masked images into refined CIF files by coupling template matching with unsupervised atom localization; the resulting mask-aware CrystalForge (EM2CIF) procedure constrains reciprocal-space search and maintains a lattice-vector RMSD of 0.09 Å on STEM2Mat. MatProphet adopts a mixture-of-experts design in which a lightweight gate-attention network orchestrates multiple interatomic predictors, improving out-of-distribution accuracy with only 2% of the training data. ScholarSeeker uses retrieval-augmented reasoning to translate structural motifs into synthesis or application hypotheses with verifiable citations, increasing retrieval costs but ensuring interpretability and provenance. Together, these agents form an extensible end-to-end spine that can absorb new vision, language, or simulation tools as they emerge. Across all four modules, we selected configurations that deliver near-optimal accuracy and robustness; however, this fidelity does introduce extra compute in stages such as CrystalForge (EM2CIF) search and multimodel property prediction.
Building a general-purpose agent for such diversity requires an architecture that does not hardcode task-specific workflows. EMSeek’s generalist performance derives from three methodological advances. First, a multimodal tool-selection layer queries a unified registry of software, databases, and models, dynamically routing each subtask to the most pertinent resource in real time. Second, code is elevated to the universal action interface, enabling the system to weave conditional logic, parallel branches and iterative refinement into its plans without hardcoding domain workflows. Third, an adaptive planning strategy grounds the initial task graph in materials knowledge and then continually revises it as intermediate results arrive, yielding context-aware behavior that resists domain drift. In concert, these innovations automate the full materials characterization pipeline, thereby accelerating hypothesis generation and closing the last mile between EM observation and materials discovery.
EMSeek enables one-for-all atom segmentation with one-click guidance
High-resolution EM routinely captures tens of thousands of atomic columns per micrograph, but manual or prompt-based tools rarely identify all symmetrically equivalent sites in a single pass. EMSeek introduces a one-for-all atom-segmentation mode: With a single click on a representative site, the model immediately retrieves every crystallographically similar column across the field of view. This capability accelerates lattice mapping, defect censuses, and precursor model construction while eliminating operator bias that would otherwise compromise downstream structure refinement and property calculations.
To train and evaluate SegMentor, we curated a benchmark that mirrors the breadth and difficulty of contemporary EM analysis. The corpus (Fig. 2D) contains thousands of pixel-annotated micrographs drawn from 20 material systems and is organized into five canonical tasks: atomic-column localization (29, 30), point-defect tagging, nanoparticle outlining (31, 32), irradiation-induced defect counting (30), and single-atom recognition (29). Materials were chosen according to three criteria: (i) relevance to current challenges in catalysis, energy storage, and semiconductor reliability; (ii) representative structural complexity spanning high-symmetry crystals to heavily defective or low-contrast lattices; and (iii) imaging diversity, with variations in acceleration voltage, dose, and detector modality that stress-test model robustness. Detailed dataset statistics, including per-task/material image counts and material breakdowns, are provided in table S1.
Fig. 2. One-click reference-patch framework for universal EM segmentation.
(A) Architecture of Ref-UNet: A user-selected reference patch from the EM image is encoded and injected via cross-attention into each encoder-decoder stage (P1 to P5) to guide mask prediction. (B) Task-level performance across five segmentation tasks—atom column identification, atom defect detection, irradiated alloy defect detection, nanoparticle recognition, and single-atom catalyst identification. (C) One-click, one-for-all atom segmentation. Left: All targets are segmented from a single user-defined patch. Right: In pristine MoS2, distinct masks are generated for Mo and S atomic columns. (D) Size of the curated multitask EM dataset spanning 20 material systems and five task types. (E) Material-wise performance across the 20 materials.
The resulting multitask dataset was used to train Ref-UNet (Fig. 2A), a U-Net (33, 34) backbone equipped with a lightweight reference encoder that injects contextual priors from a single exemplar patch. To reduce sensitivity to reference-patch quality or location, we sample patches at random positions and contrasts during training and additionally match them against the full image, thereby covering most spatial and illumination variations. Empirically, performance varies only marginally across different atomic sites and brightness levels, although users can further stabilize outputs by choosing a high-contrast, defect-free region away from image borders. Ref-UNet substantially outperforms the Segment Anything Model on every individual task and in its overall performance (Fig. 2, B and E). With 259 GFLOPs and 28 million (M) parameters, less than half the compute and roughly one-third the size of SAM 2 Hiera-B+ (560 GFLOPs and 81 M parameters), Ref-UNet runs twice as fast on a single GPU, enabling real-time interactive feedback within the EMSeek workflow.
We wrap Ref-UNet in an EMSeek Segmentation Agent that couples real-time inference with an intuitive EM viewer. As shown in Fig. 2C, each user click is translated into reference tensors, dispatched to the model, and returned as an overlay within one video frame, allowing scientists to iteratively refine masks while navigating tilt series or in situ movies. The agent exports pixel-perfect region-of-interest crops, statistical descriptors, and 3D atomic coordinates, which feed directly into automated CIF building, phase-fraction analysis, and subsequent property prediction modules. Figure 3 highlights the model’s resilience: Ref-UNet enforces periodicity and sharp boundaries (Fig. 3A), delineates irregular particles while suppressing background false positives (Fig. 3B), recovers faint point defects with fewer spurious detections (Fig. 3C), and isolates single-atom sites from noise with higher counting fidelity (Fig. 3D). By collapsing minutes of manual tracing into seconds of guided interaction, the agent transforms segmentation from a bottleneck into a seamless prelude to high-throughput materials discovery. To accommodate entirely new material classes or unseen structural motifs, the agent also supports rapid few-shot fine-tuning via transfer learning, provided users supply a small set of clear segmentation annotations. In addition to single-click operation, SegMentor supports a multireference mode in which multiple exemplar patches for the same target class are aggregated within a single run, further improving robustness to variations in size and contrast. An automatic reference-selection strategy, based on an initial U-Net prediction, is available for typical single-phase cases (see the Supplementary Materials).
Fig. 3. Comparative segmentation results of SAM and Ref-UNet across diverse microstructures.
(A) Atom column identification in DyScO3. (B) Nanoparticle recognition in PtSn@Al2O3. (C) Atom defect detection in WSe2. (D) Single-atom catalyst identification (Pt@Al2O3). For each panel, the SAM (left) and Ref-UNet (right) outputs are overlaid on the micrographs; insets show zoomed regions. Ref-UNet produces tighter boundaries and fewer false positives, particularly in dense or low-contrast areas.
EMSeek bridges EM and crystallography via one-click CIF generation
EM images are indispensable for resolving atomic structure, yet converting a noisy 2D projection into a reliable crystallographic model is fragile and labor intensive. Classical workflows that combine global template matching with manual indexing are highly sensitive to contrast drift, stage or scan instability, hydrocarbon contamination, and partial occlusion of Bragg disks. In practice, they succeed only when the correct lattice already resides in a reference library, which limits discovery settings. Recent learning-based tools such as AtomAI (24) and AutoMat (25) automate parts of the pipeline; however, because they act directly on raw pixels, their accuracy degrades under realistic noise, and they generalize poorly to materials that were not represented in training, particularly when imaging conditions shift or the structure is genuinely novel.
EMSeek avoids these failure modes with a segment-then-reconstruct strategy that isolates atomic signal before any crystallographic reasoning (Fig. 4A). A Ref-UNet first produces crystal-only masks that suppress background and scan-drift artifacts; the CrystalForge (EM2CIF) module then performs template matching and lattice refinement inside the masked regions only. Focusing the search on atoms rather than full-field pixels reduces the effective hypothesis space and removes many spurious correlations introduced by contamination. As a result, EMSeek delivers consistently higher agreement with ground truth than pixel-level baselines: S.S. (defined as a composite measure combining projected lattice RMSD and structure-matching success rate) exceeds 90% across all three difficulty tiers, outperforming both AtomAI and AutoMat in every tier (Fig. 4C, left). The improvement is most pronounced in noisy scenes where small changes in contrast or residual carbon films cause global matchers to drift.
Fig. 4. End-to-end agentic workflow linking EM images to materials knowledge.
(A) CrystalForge (EM2CIF) pipeline. Starting from EM images, atom masks are extracted and matched to candidates from a curated database or a generative library (MatterGen, Microsoft 2025). Retrieval, generative, and hybrid strategies reconstruct crystal structures. (B) MatProphet agent. The reconstructed CIFs are passed into a hub of interatomic models, whose outputs are ensembled by the MatProphet agent to predict materials properties such as energy per atom. (C) Performance evaluation of CrystalForge (EM2CIF). S.S., RMSD, and lattice error (ε) are benchmarked across tiers, comparing EMSeek with AtomAI, AutoMat, retrieval, and hybrid modes. (D) Evaluation of property prediction. Prediction errors are assessed on three out-of-distribution datasets, comparing MatProphet outputs with baseline models. (E) Uncertainty analysis: Absolute error of MatProphet Agent correlates strongly with indicators such as SD, entropy, and mean absolute deviation, demonstrating useful calibration. MAD, mean absolute deviation.
Beyond robustness, EMSeek adapts to materials outside curated catalogs. We pair database retrieval with a generative candidate library based on MatterGen, which can synthesize chemically plausible crystal hypotheses for compositions and symmetries not present in the database. EMSeek matches these generated candidates to the experimental masks and ranks them jointly with retrieved entries. This hybrid retrieval-plus-generation mode consistently improves RMSD and lattice discrepancy (ε_lattice) compared with retrieval-only or generation-only settings, with the largest gains appearing in tiers that include occlusion and acquisition shifts (Fig. 4C, middle and right). In practice, this extends reconstruction beyond any single library and enables robust CIF recovery for previously unseen chemistries while retaining the precision of database matches when they exist. EMSeek therefore supports reliable structure reconstruction in routine conditions and in exploratory regimes where noise is high, imaging varies, and the true lattice may never have been seen before.
EMSeek coordinates an agentic model ensemble for rapid property prediction
Transforming a newly reconstructed CIF into quantitative, first-principles–grade properties faces a familiar trade-off: DFT is accurate but slow, whereas individual surrogate models are fast yet fragile outside their training domains. Because no single machine-learned interatomic potential is uniformly reliable across chemistries, bonding motifs, or state points, we pair EMSeek’s reconstructed CIFs with MatProphet, a lightweight mixture-of-experts metamodel that freezes a panel of pretrained predictors [UMA (35), ORB v3 (36), MACE Foundation (37), MatterSim (38); optional domain experts] and fits only a small attention-gated affine calibrator (GatedAffineMoE; see Materials and Methods). The gate is tuned in a few-shot regime (≈2% labeled structures in our tests), enabling rapid domain adaptation while preserving each expert’s embedded physics priors.
For each structure, MatProphet fuses expert forward passes to produce consensus estimates of total energy, per-atom forces, and the stress tensor. The same component-wise scheme extends to additional scalar, vector, or tensor targets (for example, elastic moduli from finite-strain fits, defect formation energies, or migration barriers) whenever an expert provides the requisite channel. Ensemble dispersion together with learned gate weights yields an intrinsic uncertainty signal that highlights extrapolations. We compute both aleatoric and epistemic signals: Ensemble dispersion (variance across experts) and gate entropy are converted into calibrated confidence intervals that are written into the EMSeek report and passed downstream. As shown in Fig. 4E, these uncertainty indicators (SD, entropy, and mean absolute deviation) correlate strongly with absolute error, confirming their utility as reliability measures. Across three evaluation suites [JARVIS-C2DB (39), LeMat-Bulk (40), and Matbench-Discovery (41); Fig. 4D], MatProphet matches or exceeds the best individual expert and lowers overall root mean square error. To further validate performance and extract mechanistic insight, we developed ScholarSeeker (Fig. 5A), an autonomous agent built on a deep-research architecture that continuously searches, organizes, and cross-checks the literature. EMSeek uses ScholarSeeker as an evidence layer to validate or contextualize MatProphet’s outputs when uncertainty is elevated. If ScholarSeeker retrieves a citable numeric property for the exact compound (e.g., DFT total energies reported in peer-reviewed tables), then Maestro prioritizes the evidence-backed value; otherwise, the system returns MatProphet as a model-only estimate with explicit ensemble dispersion. This policy prevents unsupported citations while still enabling quantitative screening under evidence gaps. Representative examples of both interaction modes are provided in the Supplementary Materials (see Supplementary Materials, section E, and the accompanying example figures). For a given query, the agent iteratively cycles through search, information extraction, and reasoning and returns an evidence-based synthesis of current strategies. Together, these modeling and knowledge-retrieval capabilities enable uncertainty-aware decisions to be made immediately after structure reconstruction.
Fig. 5. ScholarSeeker agent for literature-grounded reasoning.
(A) The agent integrates literature search, retrieval-augmented generation (RAG), evidence extraction, and reasoning analysis, iteratively refining queries through an LLM-based loop. (B) On the Metallurgy-QA dataset, ScholarSeeker achieves higher correctness and completeness with lower hallucination and unknown rates compared with LLM and LLM + WebSearch. (C) For a query on strengthening Mg-Al-Mn cast alloys (e.g., AM100A), ScholarSeeker returns evidence-grounded recommendations on heat treatment and alloying strategies, citing domain literature (71).
Linking rapid micrograph-to-CIF reconstruction with calibrated, uncertainty-aware property prediction shortens the observe-model-screen loop: experimental structures can be triaged for stability, mechanical response, or defect energetics before launching large density-functional campaigns. Ensemble uncertainty also feeds back to rectify EMSeek: Large disagreements or unphysical outputs trigger automatic checks on segmentation masks, lattice indexing, or alternative reference patches, providing a built-in quality-assurance path that improves both structure extraction and downstream decision-making. These checks are implemented as controller-triggered reinvocations with adjusted settings while preserving prior outputs for traceability. This integrated, auditable workflow lowers the barrier to iterative cycling between microscopy, simulation, and experiment in data-limited materials discovery.
EMSeek generalizes to real-world materials tasks across diverse subfields
Materials questions posed to EM data are highly heterogeneous, ranging from atomic-scale defect inventories in layered crystals to spatial statistics of irregular features. Historically, automated solutions have depended on bespoke scripts, ad hoc data formats, and extensive manual curation. Low signal-to-noise ratios, overlap between defect and strain contrast, beam drift, and sparse labels have hindered fully automatic pipelines. Even with frameworks such as AtomAI, analysts often spend weeks retraining models for each new dataset. EMSeek addresses this fragmentation by turning user requests into natural-language prompts to a common agentic backbone. A single reference–guided segmentation model produces masks that downstream modules reuse, enabling consistent and few-shot adaptation across tasks.
Benchmarking with the latest configuration shows that this shared backbone maintains performance across image sets with only light prompt calibration. Figure 6A demonstrates lattice retrieval in a scanning transmission electron microscopy (STEM) image of monolayer-like MoS2 (42, 43): SegMentor segments Mo and S atomic columns, the mask is passed to CrystalForge (EM2CIF) to retrieve a MoS2-like lattice, and MatProphet predicts materials properties such as energy per atom. Figure 6B demonstrates nanoparticle analytics on an STEM image of PtSn nanoparticles supported on amorphous Al2O3 (44): SegMentor detects ≈73 particles with near-circular projections, and AnalyzerHub produces a size histogram indicating a moderately polydisperse ensemble. Masks, per-particle statistics, and the histogram are provenance tracked in the report. Results are produced in seconds per micrograph, and every mask, measurement, and plot is fully provenance tracked.
Fig. 6. Automated analysis based on EMSeek.
(A) The agent segments atomic columns from the raw STEM image, applies masks, and retrieves a MoS2-like lattice consistent with a quasi-2D monolayer model. Structural validation confirms lattice geometry, interlayer coupling, and spin-orbit coupling effects in agreement with known MoS2 physics in transition metal dichalcogenides (TMDCs) (42, 43, 72). (B) The agent analyzes an EM image of PtSn nanoparticles dispersed on an amorphous Al2O3 support. Segmentation masks identify ~73 particles with near-circular projections. Histogram-based particle size analysis reveals a moderately polydisperse distribution, highlighting the role of medium-sized particles in balancing catalytic activity and stability in supported PtSn systems.
Beyond imaging, EMSeek integrates literature-grounded reasoning for scientific question and answer. Figure 5 shows that incorporating domain papers through the ScholarSeeker agent yields more accurate, citation-backed answers than a base LLM or web search on Metallurgy-QA (45). Notably, naive web augmentation can degrade performance because web-scale retrieval is heterogeneous and often returns domain-adjacent but task-agnostic passages, which can introduce evidence task mismatch and method substitution in rule-constrained metallurgy questions. For completeness, detailed qualitative examples and analysis are provided in the Supplementary Materials, section G, and the associated case study figures. In contrast, by retrieving and extracting evidence from peer-reviewed sources, the agent reduces hallucinations and improves correctness and completeness. For example, when asked how to improve the strength and toughness of Mg-Al-Mn alloys, it retrieves standard heat-treatment protocols from authoritative references and returns a precise recommendation with citations. This coupling of microscopy analysis with literature retrieval bridges raw experimental data and validated knowledge, enabling evidence-based decisions across materials subfields.
EMSeek compresses EM analysis from weeks to minutes
Manual postacquisition workflows remain the chief bottleneck in EM-driven materials research. A recent in situ TEM study required three experts almost 20 weeks to fully annotate 1200 frames of irradiation defects (45), while routine static images still demand minutes to an hour per micrograph for reliable defect counts (46). Atomic-resolution lattice interpretation or grain mapping likewise consumes hours of specialist effort and is prone to analyst bias (47, 48). On the same data types, EMSeek completes reference-guided segmentation, mask-aware lattice reconstruction, property prediction, literature retrieval, validation, and report assembly in 146 ± 18 s on four A100 GPUs. This throughput reduces effective turnaround by two to three orders of magnitude, transforming EM from a retrospective diagnostic to a near–real-time driver of hypothesis testing and process optimization.
The workflow mirrors expert practice but executes it programmatically with minimal handoffs. After a user supplies a reference click, EMSeek chains segmentation, mask-aware lattice reconstruction, property estimation through the MatProphet MoE, literature checks via ScholarSeeker, and automated physical-consistency tests and then assembles figures, tables, and narrative reports. Modules stream results forward and cache intermediates so that batches of micrographs run in parallel and previously processed data can be revisited without recomputation. To scale to tens of thousands of images, the framework supports multi-GPU and distributed deployment, allowing concurrent processing of large image queues and asynchronous report generation. Every mask, parameter, script, and external source is provenance tracked, replacing ad hoc note-taking and enabling rapid review and reproducible comparison across experiments. Each agent’s computation graph and intermediate artifacts are exposed in the report, so domain experts can quickly audit masks, lattice fits, and property estimates; flag errors; and trigger targeted reruns without repeating the full pipeline. Together, these design choices account for the two– to three–order-of-magnitude reduction in analysis time documented above.
DISCUSSION
EMSeek introduces a paradigm for EM-driven materials discovery, positioning general-purpose AI agents as routine partners in structural characterization. By coupling deep-learning segmentation with literature-aware reasoning in a single multiagent framework, the platform generalizes from Å-scale atom columns to nanoscale nanoparticles and converts raw micrographs into CIF files, property predictions, and cited design suggestions—all without manual coding or task-specific tuning. This breadth mirrors the role that AI models now play in text-centric domains and signals a comparable acceleration for imaging-based materials research.
Automating what were once disjointed, labor-intensive steps free scientists to focus on hypothesis generation and experimental strategy. A novice user can upload an EM image, ask in plain language for defect statistics or elastic-modulus estimates, and receive pixel-accurate masks, lattice reconstructions, model-ensemble predictions and literature context in minutes. EMSeek could collapse days of image annotation and density-functional calculations into an interactive session, while in industrial quality control, it offers reproducible, audit-ready reports that trace every intermediate mask, script, and database call (49–51).
Several factors still constrain EMSeek and motivate ongoing development. Very low-dose, heavily contaminated, or high–dynamic-range micrographs can still degrade segmentation accuracy. To address these issues, we are integrating a self-supervised contrast normalizer that adapts the encoder to extreme signal-to-noise conditions in real time. ScholarSeeker’s performance is likewise bounded by the scope and timeliness of its text corpus. The next release will therefore ingest daily feeds from major preprint servers, extend coverage to conference proceedings and patents, and prioritize sources through citation-graph expansion to keep the literature layer closely aligned with current practice. Last, EMSeek is engineered so that the tool layer can evolve with the microscopy and materials ecosystem: Although the current catalog covers only a finite set of external and in-house routines, each is exposed as a tool with a stable key, a simple JSON-style input schema, and structured outputs, all orchestrated by the LLM-based controller, so that emerging analysis methods can be incorporated by adding new tool entries and thin wrappers without modifying the core framework. From a system perspective, the computational speedups provided by intermediate caching and batch execution are balanced by explicit limits on cache size and lifetime, lightweight modes that retain only essential outputs and compression and deduplication of image artifacts.
Future work will address these gaps in three directions. First, we will implement adaptive learning and real-time feedback loops that monitor signal-to-noise ratios, drift, and contrast shifts on the fly; user edits or high-uncertainty flags will be folded back as incremental fine-tuning data, and the agents will automatically resample reference patches, retune thresholds, or expand the retrieval library to maintain accuracy under highly variable imaging conditions. Second, extending the pipeline to 3D and in situ EM data would open avenues in operando catalysis, battery cycling, and deformation studies. Third, embedding uncertainty quantification and richer interpretability dashboards will increase user confidence and guide decision-making when predictions are deployed in costly synthesis campaigns. By unifying quantitative image analysis with qualitative knowledge extraction, EMSeek shortens the path from observation to insight and lays the groundwork for an AI-first microscope workflow. As the framework scales to additional data types and continuously refines its models through these feedback mechanisms, it is positioned to deliver more robust, transparent, and scalable materials discovery.
MATERIALS AND METHODS
EMSeek agentic platform
EMSeek operates inside a containerized workspace that ships with every resource its agents require: a Ref-UNet trained on EM images, a library of simulated diffraction patterns for template matching, multiple checkpoints for property inference, and a local mirror of crystallographic and materials databases. All artifacts—raw micrographs, intermediate masks, CIF files, and Python objects—reside in a shared key-value store whose entries are versioned and time stamped, allowing concurrent reads and writes without race conditions. Uniform API wrappers expose core functions, so that code emitted by language models interacts with the environment exactly as native Python would. To balance planning latency with numerical throughput, we pin a dedicated subset of accelerators to the LLM controller, while the remaining devices run segmentation and property-prediction kernels. This isolation avoids resource contention but can limit throughput when batch sizes spike. We expose configurable resource policies, so users can trade reasoning latency for bulk inference speed across workstations, on-prem clusters, or cloud environments.
High-level reasoning relies on instruction-tuned GPT-4 and DeepSeek class models that plan workflows and conduct dialog, supported by a distilled low-latency variant for rapid tool retrieval and a multimodal encoder-decoder adapted from GPT-4o when visual context is needed. Vision inference is handled by the Ref-UNet backbone, whereas property prediction fuses outputs from UMA (35), ORB v3 (36), MACE Foundation (37), MatterSim (38), and optional domain experts. We chose GPT-based and DeepSeek models because their long-context reasoning, tool-use proficiency, and safety alignment currently outperform available alternatives in our tasks. Nevertheless, EMSeek exposes a modular LLM interface so the controller can be swapped for advanced open-source or multimodal transformers as they mature, enabling users to balance cost, privacy, or domain specialization against raw capability. In future deployments, richer cross-modal planners or domain-tuned open weights could further enhance agentic reasoning without altering the surrounding pipeline.
Agents communicate through the shared store and a lightweight message bus, exchanging natural-language action requests accompanied by JSON schemas that specify expected outputs. All action requests are timestamped and pushed to a queue. Within this queue, the agents run an observe-think-act loop: They atomically pop the latest state, reason with their language model, execute code, and commit results as new, version-incremented keys. Maestro Agent decomposes a user query into subtasks, allocates them to specialists, tags each subtask with a unique query ID so that multiple user sessions can interleave safely in the same queue, and reassigns work whenever confidence scores fall below threshold; Guardian Agent enforces domain constraints—physical plausibility of lattice parameters, unit consistency, and citation validity—and triggers rollbacks when violations arise; Scribe Agent assembles final artifacts into a provenance-linked report, embedding figures, tables, and executable code cells so that every result can be audited or reproduced. Detailed prompt templates, hyperparameters, and fallback strategies are provided in the Supplementary Materials.
Reference-guided segmentation benchmark and validation
To benchmark reference-guided segmentation, we curated a dataset spanning more than 20 material systems—perovskites, high-entropy alloys, van der Waals heterostructures, single-atom catalysts, and others—in total 20 distinct compositions and structures spanning 11 crystallographic space groups (29–31, 52–56). Five targets received pixel-accurate, consensus annotations: atomic columns, point defects, single-atom adsorbates, nanoparticle outlines, and irradiation-induced dislocation networks. As illustrated in Fig. 2D, a substantial portion of the samples focus on atomic columns, underscoring the pivotal role of atomic-scale resolution in contemporary materials research. In addition, the inclusion of broader categories—such as nanoparticles and 2D materials—ensures that EMSeek’s segmentation and analysis agents remain versatile and generalizable across diverse structural and compositional domains. The dataset combines publicly available EM images with our own collected scans; full composition, provenance, and licensing details are provided in the Supplementary Materials, and the complete dataset (images, masks, and metadata schema) is openly released. We also expose a submission portal to enable community contributions, allowing the benchmark to expand and further strengthen model generalization and validation.
SegMentor Agent
At the heart of SegMentor is Ref-UNet (Fig. 2A), a lightweight U-Net whose encoder blocks are replaced by a visual backbone and whose skip connections carry not only feature maps but also a learned embedding of a user-selected reference patch. During the forward pass, the patch is tokenized, position encoded, and injected into every encoder stage through cross-attention layers that reweight channel responses according to patch similarity; the resulting context vector travels with the up-sampling path, steering pixel-wise predictions toward features that match the reference while suppressing distractors. The network is trained end to end on random paired images and masks that span 21 material systems, using a composite Dice-plus-binary-cross-entropy loss and aggressive augmentation (rotations, flips, Poisson noise, and contrast jitter) to harden it against beam-current drift and detector artifacts. At inference time, SegMentor ingests the reference patch coordinates from the web interface, runs one forward pass, applies a small-object filter, and writes the mask to shared memory for downstream agents. The final prediction is
| (1) |
where are image features at scale , , and is a patch-informed channel gate computed from cross-attention between queries from the patch and keys from the image (sigmoid of elementwise with channel averaging). Here, is the dense U-Net++ decoder, and σ is the sigmoid.
CrystalForge (EM2CIF) benchmark and validation
To ensure strict comparability with the published AutoMat study, we evaluated CrystalForge (EM2CIF) on the STEM2Mat benchmark, which contains 450 aberration-corrected STEM micrographs partitioned into three tiers: “easy” images with high signal-to-noise ratio and minimal contamination, “medium” images with moderate drift or partial staining, and “hard” images with low contrast, heavy residue, or severe scan artifacts. Unlike AutoMat, which performs template matching over the entire field of view, CrystalForge (EM2CIF) uses the binary lattice mask produced by SegMentor to restrict retrieval to that footprint, and atom localization is initialized at mask centroids rather than uniformly across the image. To align CIF candidates with the experimental micrograph, we first expand each CIF into the smallest supercell that preserves local periodicity and render a clean projection . We then perform small-template to large-image matching by sliding a window over the masked region of the experimental image and comparing band-limited fast Fourier transform magnitude features . The window score is the cosine similarity
| (2) |
computed after removing the dc component and retaining a fixed frequency band to reduce illumination and noise effects. We search a small grid of rotations and scales for robustness, select the best location to initialize refinement, and aggregate evidence across the SegMentor footprint using a weighted average of the sliding-window scores. The reconstructed CIFs were benchmarked using three metrics: S.S., RMSD of lattice vectors, and relative lattice error (ε). S.S. was directly compared with the baseline systems AtomAI, AutoMat, and EMSeek, while RMSD and lattice error were separately assessed for retrieval-only, generative-only, and hybrid strategies. This evaluation showed that CrystalForge (EM2CIF) achieves higher S.S. than prior approaches on lower-tier images and that hybrid strategies provide balanced improvements in RMSD and lattice error relative to retrieval or generative methods alone. As mask quality deteriorates on the hard tier, CrystalForge (EM2CIF) performance decreases proportionally, with higher RMSD and ε\varepsilonε and lower success rates, indicating that noise-induced mask errors remain the dominant failure mode. These results highlight the advantages of mask-aware search combined with generative augmentation and emphasize the importance of adaptive preprocessing and dynamic mask refinement under variable imaging conditions.
Property prediction benchmark and validation
We assessed whether EMSeek can convert newly reconstructed CIFs into reliable materials properties without task-specific retraining by evaluating MatProphet on three out-of-distribution suites: JARVIS-C2DB (39), LeMat-Bulk (40), and Matbench-Discovery (41). MatProphet holds four pretrained interatomic predictors fixed: UMA (35), ORB v3 (36), MACE Foundation (37), and MatterSim (38). Each CIF is standardized and broadcast to all experts within a common container. Returned quantities are converted to consistent units (total energy in electron volts per atom, forces in electron volts per angstrom, and stress in gigapascals in Voigt six-component form). Channels not provided by an expert are masked during gating. For GatedAffineMoE, expert predictions for a batch form a tensor of shape batch × experts. We compute attention-based gates by projecting the expert stack with a learnable query to obtain = , comparing to a learned key matrix via scaled dot product to produce logits, refining them with a two-layer multi-layer perceptron (MLP) using rectified linear unit (ReLU) activations, and applying softmax to yield normalized weights . Each expert’s output is then affine calibrated with a learnable per-expert scale and bias . The final prediction adds a small residual of the mean uncalibrated experts to stabilize training
| (3) |
where α = 0.1 by default. Parameters of the MoE head are fitted on ~2% labeled structures from each benchmark (stratified by energy range) using Smooth L1 loss, cosine learning-rate decay, and early stopping on a held-out validation split; all backbone experts remain frozen. The remaining entries constitute the test sets.
ScholarSeeker Agent and validation
To provide EMSeek with up-to-date scientific context, we constructed a literature corpus comprising several thousand peer-reviewed articles from the past 5 years. Each source was parsed from PDF into a structured representation including metadata, section headers, paragraph text, figure captions, and references (57–62). Text spans were segmented at the sentence level, embedded with the OpenAI text-embedding-3-small model, and indexed in a vector database together with bibliographic metadata such as DOI, journal, year, and section label. ScholarSeeker operates in a three-stage loop: Literature search retrieves candidate passages through dense similarity queries; evidence extraction ranks and filters sentences, recording full provenance via the Guardian agent (DOI and sentence offset); and reasoning analysis organizes the evidence into structured arguments, with the Scribe agent injecting citations and supporting text into the user-facing report. This design enforces citation integrity and evidence traceability. For rapidly evolving subfields, we track recall of articles published within the last 6 months, schedule periodic corpus refreshes, and maintain an automated alert pipeline that flags citation gaps when user queries reference unseen DOIs. Community submission hooks allow users to propose missing literature, which is validated and embedded automatically. On the Metallurgy-QA benchmark (63), ScholarSeeker demonstrates improved correctness and completeness with reduced hallucination and unknown rates compared with LLM baselines.
AnalyzerHub Agent
To give EMSeek immediate access to state-of-the-art analytics, we integrated leading open-source packages for EM and materials modeling, including HyperSpy (64), Py4DSTEM (65), Atomap (66), and so on (67–70). AnalyzerHub, a retrieval augmented language model, selects, sequences, and executes tools from the registry. Given a user query or a request from another agent, it performs dense embedding search over the registry descriptors, ranks candidate tools by capability and estimated computational cost, emits executable Python snippets, and streams results back to shared memory for immediate reuse by MatProphet or Scribe. This dynamic, sequential chaining avoids launching all tools at once, reducing peak memory and GPU utilization, although it can introduce cumulative latency when workflows span many modules. To mitigate this, we schedule lightweight tools earlier in the chain, cache reusable intermediates, and allow optional parallel execution when independent branches are detected. Users may also explicitly pin or skip tools via natural-language directives. All intermediate data, parameters, and code calls are logged to preserve full provenance.
Acknowledgments
Funding:
This project is supported by the Eric and Wendy Schmidt AI in Science Postdoctoral Fellowship, a program of Schmidt Sciences LLC. The material is based upon work partially supported by the National Science Foundation (NSF) through an NSF Research Traineeship (NRT) project under grant no. 2345579.
Author contributions:
F.Y. conceived the study. W.Y. collected and annotated the data. G.C. developed the models and analyzed the results. G.C., W.Y., and F.Y. contributed to writing the manuscript, and all authors reviewed and approved the final version.
Competing interests:
The authors declare that they have no competing interests.
Data, code, and materials availability:
This study did not generate new materials. The multitask EM dataset and source code are available on Zenodo at https://doi.org/10.5281/zenodo.18340508. The source code is also publicly available at https://github.com/PEESEgroup/EMSeek. All data and code needed to evaluate and reproduce the results in the paper are present in the paper and/or the Supplementary Materials.
Supplementary Materials
The PDF file includes:
Supplementary Sections A to I
Figs. S1 to S22
Tables S1 and S2
Legend for movie S1
Other Supplementary Material for this manuscript includes the following:
Movie S1
REFERENCES
- 1.Zhu Y., Ciston J., Zheng B., Miao X., Czarnik C., Pan Y., Sougrat R., Lai Z., Hsiung C.-E., Yao K., Pinnau I., Pan M., Han Y., Unravelling surface and interfacial structures of a metal–organic framework by transmission electron microscopy. Nat. Mater. 16, 532–536 (2017). [DOI] [PubMed] [Google Scholar]
- 2.Zhao C., Jiang Z., Liu Y., Zhou Y., Yin P., Ke Y., Deng H., Molecular compartments created in metal–organic frameworks for efficient visible-light-driven CO2 overall all conversion. J. Am. Chem. Soc. 144, 23560–23571 (2022). [DOI] [PubMed] [Google Scholar]
- 3.Yan X., Liu C., Gadre C. A., Gu L., Aoki T., Lovejoy T. C., Dellby N., Krivanek O. L., Schlom D. G., Wu R., Pan X., Single-defect phonons imaged by electron microscopy. Nature 589, 65–69 (2021). [DOI] [PubMed] [Google Scholar]
- 4.Gadre C. A., Yan X., Song Q., Li J., Gu L., Huyan H., Aoki T., Lee S.-W., Chen G., Wu R., Pan X., Nanoscale imaging of phonon dynamics by electron microscopy. Nature 606, 292–297 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Weirich T. E., Ramlau R., Simon A., Hovmöller S., Zou X., A crystal structure determined with 0.02 Å accuracy by electron microscopy. Nature 382, 144–146 (1996). [Google Scholar]
- 6.Hashimoto H., Yotsumoto H., Spacing anomaly in electron microscope images of crystal lattices. Nature 183, 1001–1002 (1959). [Google Scholar]
- 7.Chadderton L. T., Electron microscopy of crystal lattices: An anomalous effect. Nature 189, 564–565 (1961). [Google Scholar]
- 8.Sun B., Lu W., Gault B., Ding R., Makineni S. K., Wan D., Wu C.-H., Chen H., Ponge D., Raabe D., Chemical heterogeneity enhances hydrogen resistance in high-strength steels. Nat. Mater. 20, 1629–1634 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Choe J., Lee Y., Park J., Kim Y., Kim C. U., Kim K., Direct imaging of structural disordering and heterogeneous dynamics of fullerene molecular liquid. Nat. Commun. 10, 4395 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Jiao Y., Qiu Y., Zhang L., Liu W.-G., Mao H., Chen H., Feng Y., Cai K., Shen D., Song B., Chen X.-Y., Li X., Zhao X., Young R. M., Stern C. L., Wasielewski M. R., Astumian R. D., Goddard W. A. III, Stoddart J. F., Electron-catalysed molecular recognition. Nature 603, 265–270 (2022). [DOI] [PubMed] [Google Scholar]
- 11.Bai X., McMullan G., Scheres S. H. W., How cryo-EM is revolutionizing structural biology. Trends Biochem. Sci. 40, 49–57 (2015). [DOI] [PubMed] [Google Scholar]
- 12.G. Chen, S. Dong, Y. Shu, G. Zhang, J. Sesay, B. Karlsson, J. Fu, Y. Shi, “AutoAgents: A framework for automatic agent generation,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (International Joint Conferences on Artificial Intelligence Organization, 2024), pp. 22–30. [Google Scholar]
- 13.T. Guo, X. Chen, Y. Wang, R. Chang, S. Pei, N. V. Chawla, O. Wiest, X. Zhang, “Large language model based multi-agents: A survey of progress and challenges,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (International Joint Conferences on Artificial Intelligence Organization, 2024), pp. 8048–8057. [Google Scholar]
- 14.MacLeod B. P., Parlane F. G. L., Morrissey T. D., Häse F., Roch L. M., Dettelbach K. E., Moreira R., Yunker L. P. E., Rooney M. B., Deeth J. R., Lai V., Ng G. J., Situ H., Zhang R. H., Elliott M. S., Haley T. H., Dvorak D. J., Aspuru-Guzik A., Hein J. E., Berlinguette C. P., Self-driving laboratory for accelerated discovery of thin-film materials. Sci. Adv. 6, eaaz8867 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Volk A. A., Abolhasani M., Performance metrics to unleash the power of self-driving labs in chemistry and materials science. Nat. Commun. 15, 1378 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Abolhasani M., Kumacheva E., The rise of self-driving labs in chemical and materials sciences. Nat. Synth. 2, 483–492 (2023). [Google Scholar]
- 17.Szymanski N. J., Rendy B., Fei Y., Kumar R. E., He T., Milsted D., McDermott M. J., Gallant M., Cubuk E. D., Merchant A., Kim H., Jain A., Bartel C. J., Persson K., Zeng Y., Ceder G., An autonomous laboratory for the accelerated synthesis of novel materials. Nature 624, 86–91 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Ghafarollahi A., Buehler M. J., Automating alloy design and discovery with physics-aware multimodal multiagent AI. Proc. Natl. Acad. Sci. U.S.A. 122, e2414074122 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Stokes J. M., Yang K., Swanson K., Jin W., Cubillos-Ruiz A., Donghia N. M., MacNair C. R., French S., Carfrae L. A., Bloom-Ackermann Z., Tran V. M., Chiappino-Pepe A., Badran A. H., Andrews I. W., Chory E. J., Church G. M., Brown E. D., Jaakkola T. S., Barzilay R., Collins J. J., A deep learning approach to antibiotic discovery. Cell 180, 688–702.e13 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Kusne A. G., Yu H., Wu C., Zhang H., Hattrick-Simpers J., DeCost B., Sarker S., Oses C., Toher C., Curtarolo S., Davydov A. V., Agarwal R., Bendersky L. A., Li M., Mehta A., Takeuchi I., On-the-fly closed-loop materials discovery via Bayesian active learning. Nat. Commun. 11, 5966 (2020). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Burger B., Maffettone P. M., Gusev V. V., Aitchison C. M., Bai Y., Wang X., Li X., Alston B. M., Li B., Clowes R., Rankin N., Harris B., Sprick R. S., Cooper A. I., A mobile robotic chemist. Nature 583, 237–241 (2020). [DOI] [PubMed] [Google Scholar]
- 22.Coutant A., Roper K., Trejo-Banos D., Bouthinon D., Carpenter M., Grzebyta J., Santini G., Soldano H., Elati M., Ramon J., Rouveirol C., Soldatova L. N., King R. D., Closed-loop cycles of experiment design, execution, and learning accelerate systems biology model development in yeast. Proc. Natl. Acad. Sci. U.S.A. 116, 18142–18147 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.King R. D., Rowland J., Oliver S. G., Young M., Aubrey W., Byrne E., Liakata M., Markham M., Pir P., Soldatova L. N., Sparkes A., Whelan K. E., Clare A., The automation of science. Science 324, 85–89 (2009). [DOI] [PubMed] [Google Scholar]
- 24.Ziatdinov M., Ghosh A., Wong C. Y., Kalinin S. V., AtomAI framework for deep learning analysis of image and spectroscopy data in electron and scanning probe microscopy. Nat. Mach. Intell. 4, 1101–1112 (2022). [Google Scholar]
- 25.Y. Yang, Y. Tang, Y. Chen, X. Chen, J. Qiu, H. Xiong, H. Yin, Z. Luo, Y. Zhang, S. Tao, W. Li, Q. Zhang, Y. Li, W. Ouyang, B. Zhao, X. Wang, F. Wei, AutoMat: Enabling automated crystal structure reconstruction from microscopy via agentic tool use. arXiv:2505.12650 [cs.CV] (2025).
- 26.Archit A., Freckmann L., Nair S., Khalid N., Hilt P., Rajashekar V., Freitag M., Teuber C., Buckley G., Von Haaren S., Gupta S., Dengel A., Ahmed S., Pape C., Segment anything for microscopy. Nat. Methods 22, 579–591 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 27.Torquato S., Kim J., Nonlocal effective electromagnetic wave characteristics of composite media: Beyond the quasistatic regime. Phys. Rev. X 11, 021002 (2021). [Google Scholar]
- 28.Calogero G., Raciti D., Acosta-Alba P., Cristiano F., Deretzis I., Fisicaro G., Huet K., Kerdilès S., Sciuto A., La Magna A., Multiscale modeling of ultrafast melting phenomena. npj Comput. Mater. 8, 36 (2022). [Google Scholar]
- 29.Lin R., Zhang R., Wang C., Yang X.-Q., Xin H. L., TEMImageNet training library and AtomSegNet deep-learning models for high-precision atom segmentation, localization, denoising, and deblurring of atomic-resolution images. Sci. Rep. 11, 5386 (2021). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Khan A., Lee C.-H., Huang P. Y., Clark B. K., Leveraging generative adversarial networks to create realistic scanning transmission electron microscopy images. npj Comput. Mater. 9, 85 (2023). [Google Scholar]
- 31.Treder K. P., Huang C., Bell C. G., Slater T. J. A., Schuster M. E., Özkaya D., Kim J. S., Kirkland A. I., nNPipe: A neural network pipeline for automated analysis of morphologically diverse catalyst systems. npj Comput. Mater. 9, 18 (2023). [Google Scholar]
- 32.Horwath J. P., Zakharov D. N., Mégret R., Stach E. A., Understanding important features of deep learning models for segmentation of high-resolution transmission electron microscopy images. npj Comput. Mater. 6, 108 (2020). [Google Scholar]
- 33.Z. Zhou, M. M. Rahman Siddiquee, N. Tajbakhsh, J. Liang, “UNet++: A nested U-Net architecture for medical image segmentation,” in Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, vol. 11045 of Lecture Notes in Computer Science, D. Stoyanov, Z. Taylor, G. Carneiro, T. Syeda-Mahmood, A. Martel, L. Maier-Hein, J. M. R. S. Tavares, A. Bradley, J. P. Papa, V. Belagiannis, J. C. Nascimento, Z. Lu, S. Conjeti, M. Moradi, H. Greenspan, A. Madabhushi, Eds. (Springer International Publishing, 2018), pp. 3–11. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.O. Ronneberger, P. Fischer, T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015, vol. 9351 of Lecture Notes in Computer Science, N. Navab, J. Hornegger, W. M. Wells, A. F. Frangi, Eds. (Springer International Publishing, 2015), pp. 234–241. [Google Scholar]
- 35.B. M. Wood, M. Dzamba, X. Fu, M. Gao, M. Shuaibi, L. Barroso-Luque, K. Abdelmaqsoud, V. Gharakhanyan, J. R. Kitchin, D. S. Levine, K. Michel, A. Sriram, T. Cohen, A. Das, A. Rizvi, S. J. Sahoo, Z. W. Ulissi, C. L. Zitnick, UMA: A family of universal models for atoms. arXiv:2506.23971 [cs.LG] (2025).
- 36.B. Rhodes, S. Vandenhaute, V. Šimkus, J. Gin, J. Godwin, T. Duignan, M. Neumann, Orb-v3: Atomistic simulation at scale. arXiv:2504.06231 [cond-mat.mtrl-sci] (2025).
- 37.I. Batatia, P. Benner, Y. Chiang, A. M. Elena, D. P. Kovács, J. Riebesell, X. R. Advincula, M. Asta, M. Avaylon, W. J. Baldwin, F. Berger, N. Bernstein, A. Bhowmik, S. M. Blau, V. Cărare, J. P. Darby, S. De, F. D. Pia, V. L. Deringer, R. Elijošius, Z. El-Machachi, F. Falcioni, E. Fako, A. C. Ferrari, A. Genreith-Schriever, J. George, R. E. A. Goodall, C. P. Grey, P. Grigorev, S. Han, W. Handley, H. H. Heenen, K. Hermansson, C. Holm, J. Jaafar, S. Hofmann, K. S. Jakob, H. Jung, V. Kapil, A. D. Kaplan, N. Karimitari, J. R. Kermode, N. Kroupa, J. Kullgren, M. C. Kuner, D. Kuryla, G. Liepuoniute, J. T. Margraf, I.-B. Magdău, A. Michaelides, J. H. Moore, A. A. Naik, S. P. Niblett, S. W. Norwood, N. O’Neill, C. Ortner, K. A. Persson, K. Reuter, A. S. Rosen, L. L. Schaaf, C. Schran, B. X. Shi, E. Sivonxay, T. K. Stenczel, V. Svahn, C. Sutton, T. D. Swinburne, J. Tilly, C. van der Oord, E. Varga-Umbrich, T. Vegge, M. Vondrák, Y. Wang, W. C. Witt, F. Zills, G. Csányi, A foundation model for atomistic materials chemistry. arXiv:2401.00096 [physics.chem-ph] (2024).
- 38.H. Yang, C. Hu, Y. Zhou, X. Liu, Y. Shi, J. Li, G. Li, Z. Chen, S. Chen, C. Zeni, M. Horton, R. Pinsler, A. Fowler, D. Zügner, T. Xie, J. Smith, L. Sun, Q. Wang, L. Kong, C. Liu, H. Hao, Z. Lu, MatterSim: A deep learning atomistic model across elements, temperatures and pressures. arXiv:2405.04967 [cond-mat.mtrl-sci] (2024).
- 39.Haastrup S., Strange M., Pandey M., Deilmann T., Schmidt P. S., Hinsche N. F., Gjerding M. N., Torelli D., Larsen P. M., Riis-Jensen A. C., Gath J., Jacobsen K. W., Jørgen Mortensen J., Olsen T., Thygesen K. S., The Computational 2D Materials Database: High-throughput modeling and discovery of atomically thin crystals. 2D Mater. 5, 042002 (2018). [Google Scholar]
- 40.LeMaterial, LeMat-Bulk, version a36d6af, Hugging Face (2024); 10.57967/HF/3762. [DOI]
- 41.Riebesell J., Goodall R. E. A., Benner P., Chiang Y., Deng B., Ceder G., Asta M., Lee A. A., Jain A., Persson K. A., A framework to evaluate machine learning crystal stability predictions. Nat. Mach. Intell. 7, 836–847 (2025). [Google Scholar]
- 42.Marinov K., Avsar A., Watanabe K., Taniguchi T., Kis A., Resolving the spin splitting in the conduction band of monolayer MoS2. Nat. Commun. 8, 1938 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Wang Q. H., Kalantar-Zadeh K., Kis A., Coleman J. N., Strano M. S., Electronics and optoelectronics of two-dimensional transition metal dichalcogenides. Nat. Nanotechnol. 7, 699–712 (2012). [DOI] [PubMed] [Google Scholar]
- 44.Yuan W., Yao B., Tan S., He Q., Generative AI enables label-free segmentation for live analysis of supported nanoparticle catalysts. Microsc. Microanal. 30, 10.1093/mam/ozae044.211 (2024). [Google Scholar]
- 45.Sainju R., Chen W.-Y., Schaefer S., Yang Q., Ding C., Li M., Zhu Y., DefectTrack: A deep learning-based multi-object tracking algorithm for quantitative defect analysis of in-situ TEM videos in real-time. Sci. Rep. 12, 15705 (2022). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Li W., Field K. G., Morgan D., Automated defect analysis in electron microscopic images. npj Comput. Mater. 4, 36 (2018). [Google Scholar]
- 47.Huang W., Jin Y., Li Z., Yao L., Chen Y., Luo Z., Zhou S., Lin J., Liu F., Gao Z., Cheng J., Zhang L., Ouyang F., Zhang J., Wang S., Auto-resolving the atomic structure at van der Waals interfaces using a generative model. Nat. Commun. 16, 2927 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 48.Sohn W., Kim T., Moon C. W., Shin D., Park Y., Jin H., Baik H., Unsupervised learning for the automatic counting of grains in nanocrystals and image segmentation at the atomic resolution. Nanomaterials 14, 1614 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 49.Vu T.-S., Ha M.-Q., Nguyen D.-N., Nguyen V.-C., Abe Y., Tran T., Tran H., Kino H., Miyake T., Tsuda K., Dam H.-C., Towards understanding structure–property relations in materials with interpretable deep learning. npj Comput. Mater. 9, 215 (2023). [Google Scholar]
- 50.Wellawatte G. P., Schwaller P., Human interpretable structure-property relationships in chemistry using explainable machine learning and large language models. Commun. Chem. 8, 11 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Hu Z., Yan W., Data-driven modeling of process-structure-property relationships in metal additive manufacturing. npj Adv. Manuf. 1, 3 (2024). [Google Scholar]
- 52.Mitchell S., Parés F., Faust Akl D., Collins S. M., Kepaptsoglou D. M., Ramasse Q. M., Garcia-Gasulla D., Pérez-Ramírez J., López N., Automated image analysis for single-atom detection in catalytic materials by transmission electron microscopy. J. Am. Chem. Soc. 144, 8018–8029 (2022). [DOI] [PubMed] [Google Scholar]
- 53.Yuan W., Yao B., Tan S., You F., He Q., Generative learning of morphological and contrast heterogeneities for self-supervised electron micrograph segmentation. npj Comput. Mater. 11, 322 (2025). [Google Scholar]
- 54.Sytwu K., Rangel DaCosta L., Scott M. C., Generalization across experimental parameters in neural network analysis of high-resolution transmission electron microscopy datasets. Microsc. Microanal. 30, 85–95 (2024). [DOI] [PubMed] [Google Scholar]
- 55.Oktay A. B., Gurses A., Automatic detection, localization and segmentation of nano-particles with deep learning in microscopy images. Micron 120, 113–119 (2019). [DOI] [PubMed] [Google Scholar]
- 56.Jacobs R., Shen M., Liu Y., Hao W., Li X., He R., Greaves J. R. C., Wang D., Xie Z., Huang Z., Wang C., Field K. G., Morgan D., Performance and limitations of deep learning semantic segmentation of multiple defects in transmission electron micrographs. Cell Rep. Phys. Sci. 3, 100876 (2022). [Google Scholar]
- 57.S. Narayanan, J. D. Braza, R.-R. Griffiths, M. Ponnapati, A. Bou, J. M. Laurent, O. Kabeli, G. Wellawatte, S. Cox, S. G. Rodriques, A. White, “Aviary: Training language agents on challenging scientific tasks” (2025); https://openreview.net/forum?id=25Grz6oh7d.
- 58.M. D. Skarlinski, S. Cox, J. M. Laurent, J. D. Braza, M. Hinks, M. J. Hammerling, M. Ponnapati, S. G. Rodriques, A. D. White, Language agents achieve superhuman synthesis of scientific knowledge. arXiv:2409.13740 [cs.CL] (2024).
- 59.J. Lála, O. O’Donoghue, A. Shtedritski, S. Cox, S. G. Rodriques, A. D. White, PaperQA: Retrieval-Augmented Generative agent for scientific research. arXiv:2312.07559 [cs.CL] (2023).
- 60.Hoque M. N. F., Yang M., Li Z., Islam N., Pan X., Zhu K., Fan Z., Polarization and dielectric study of methylammonium lead iodide thin film to reveal its nonferroelectric nature under solar cell operating conditions. ACS Energy Lett. 1, 142–149 (2016). [Google Scholar]
- 61.Garten L. M., Moore D. T., Nanayakkara S. U., Dwaraknath S., Schulz P., Wands J., Rockett A., Newell B., Persson K. A., Trolier-McKinstry S., Ginley D. S., The existence and impact of persistent ferroelectric domains in MAPbI3. Sci. Adv. 5, eaas9311 (2019). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.W. W. Wright, “Materials science and engineering. An introduction (2nd edition),” in Polymer International, W. D. Callister Jr (John Wiley & Sons, 1993), pp. 282–283. [Google Scholar]
- 63.A. Eldeeb, Metallurgy and materials science knowledge extraction dataset (2024); https://huggingface.co/datasets/AbdulrhmanEldeeb/metallurgy-qa.
- 64.F. De La Peña, T. Ostasevicius, V. T. Fauske, P. Burdet, P. Jokubauskas, M. Nord, E. Prestat, M. Sarahan, K. E. MacArthur, D. N. Johnstone, J. Taillon, J. Caron, T. Furnival, A. Eljarrat, S. Mazzucco, V. Migunov, T. Aarholt, M. Walls, F. Winkler, B. Martineau, G. Donval, E. R. Hoglund, I. Alxneit, I. Hjorth, L. F. Zagonel, A. Garmannslund, C. Gohlke, I. Iyengar, Huang-Wei Chang, hyperspy/hyperspy: HyperSpy 1.3, version v1.3, Zenodo (2017); 10.5281/ZENODO.583693. [DOI]
- 65.Savitzky B. H., Zeltmann S. E., Hughes L. A., Brown H. G., Zhao S., Pelz P. M., Pekin T. C., Barnard E. S., Donohue J., Rangel DaCosta L., Kennedy E., Xie Y., Janish M. T., Schneider M. M., Herring P., Gopal C., Anapolsky A., Dhall R., Bustillo K. C., Ercius P., Scott M. C., Ciston J., Minor A. M., Ophus C., py4DSTEM: A software package for four-dimensional scanning transmission electron microscopy data analysis. Microsc. Microanal. 27, 712–743 (2021). [DOI] [PubMed] [Google Scholar]
- 66.Nord M., Vullum P. E., MacLaren I., Tybell T., Holmestad R., Atomap: A new software tool for the automated analysis of atomic resolution images using two-dimensional Gaussian fitting. Adv. Struct. Chem. Imaging 3, 9 (2017). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Van Der Walt S., Schönberger J. L., Nunez-Iglesias J., Boulogne F., Warner J. D., Yager N., Gouillart E., Yu T., scikit-image: Image processing in Python. PeerJ 2, e453 (2014). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Ong S. P., Richards W. D., Jain A., Hautier G., Kocher M., Cholia S., Gunter D., Chevrier V. L., Persson K. A., Ceder G., Python Materials Genomics (pymatgen): A robust, open-source python library for materials analysis. Comput. Mater. Sci. 68, 314–319 (2013). [Google Scholar]
- 69.Horton M. K., Huck P., Yang R. X., Munro J. M., Dwaraknath S., Ganose A. M., Kingsbury R. S., Wen M., Shen J. X., Mathis T. S., Kaplan A. D., Berket K., Riebesell J., George J., Rosen A. S., Spotte-Smith E. W. C., McDermott M. J., Cohen O. A., Dunn A., Kuner M. C., Rignanese G.-M., Petretto G., Waroquiers D., Griffin S. M., Neaton J. B., Chrzan D. C., Asta M., Hautier G., Cholia S., Ceder G., Ong S. P., Jain A., Persson K. A., Accelerated data-driven materials science with the Materials Project. Nat. Mater. 24, 1522–1532 (2025). [DOI] [PubMed] [Google Scholar]
- 70.Mongkhonratanachai M., Jana S., Maggard P. A., Interplay of dual-site metal disorder on the thermodynamic stability and electronic structures of Cu( i )-containing molybdovanadate and tungstovanadate compounds. Dalton Trans. 54, 17102–17111 (2025). [DOI] [PubMed] [Google Scholar]
- 71.T. V. Rajan, C. P. Sharma, A. Sharma, Heat Treatment: Principles and Techniques (PHI Learning Private Limited, New Delhi, India, 2nd ed., 2011). [Google Scholar]
- 72.Alidoust N., Bian G., Xu S.-Y., Sankar R., Neupane M., Liu C., Belopolski I., Qu D.-X., Denlinger J. D., Chou F.-C., Hasan M. Z., Observation of monolayer valence band spin-orbit effect and induced quantum well states in MoX2. Nat. Commun. 5, 4673 (2014). [DOI] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Supplementary Sections A to I
Figs. S1 to S22
Tables S1 and S2
Legend for movie S1
Movie S1
Data Availability Statement
This study did not generate new materials. The multitask EM dataset and source code are available on Zenodo at https://doi.org/10.5281/zenodo.18340508. The source code is also publicly available at https://github.com/PEESEgroup/EMSeek. All data and code needed to evaluate and reproduce the results in the paper are present in the paper and/or the Supplementary Materials.






