1. INTRODUCTION
Powder-based manufacturing, including powder metallurgy (PM) and additive manufacturing
(AM), has become a vital route for producing high-value metallic components in aerospace,
biomedical, energy, tooling, and advanced structural applications[1–
6]. These processes offer substantial flexibility in alloy design, microstructural control,
and geometrical complexity because the final product is built directly from powder
feedstock rather than from a pre-shaped bulk material. This feature gives powder-based
routes a distinct competitive advantage when conventional manufacturing is restricted
by design constraints, material yield losses, or the necessity for near-net-shape
production. At the same time, it also renders the process chain significantly more
sensitive to variations in powder state, processing history, and post-processing conditions
than traditional manufacturing routes.
The primary challenge lies in the fact that powder-based manufacturing is highly complex
and cannot be governed by a single processing step. Powder production, storage, mixing,
reuse, spreading or compaction, melting or sintering, heat treatment, hot isostatic
pressing (HIP), machining, and final inspection all systematically contribute to the
final material response. Powder descriptors such as particle size distribution, morphology,
flowability, apparent density, oxygen content, and surface condition directly affect
how powders spread, pack, absorb energy, melt, sinter, and form defects[4,
6–
8]. In powder bed fusion (PBF), for instance, powder-bed quality and recoating behavior
strongly influence melt-pool formation and the occurrence of defects, whereas powder
reuse can change particle morphology, chemistry, flow behavior, and traceability across
consecutive build cycles[9–
11]. Subsequent post-processing further modifies porosity, residual stress, microstructure,
precipitate state, and mechanical performance[22]. These variables are strongly coupled rather than sequentially independent, which
makes empirical trial-and-error optimization expensive, time-consuming, and difficult
to generalize across different materials, machines, and production environments[4,
6,
7,
10–
13].
Consequently, ML has attracted significant interest as a data-driven tool for extracting
complex relationships from powder-based manufacturing datasets. Prior studies and
reviews have demonstrated that ML can effectively support process parameter optimization,
powder-spreading assessment, in-situ monitoring, defect detection, microstructure
prediction, property prediction, quality control, and inverse design[12–
15]. Depending on the data structure, tree-based models, Gaussian process regression,
support vector machines, artificial neural networks (ANNs), convolutional neural networks
(CNNs), recurrent models, and physics-informed neural networks (PINNs) have all been
implemented to map process conditions to quality indicators or to infer processing
windows from target properties[14–
16]. These developments demonstrate that ML is no longer merely a peripheral tool in
powder manufacturing; it is rapidly becoming an integral component of the emerging
digital infrastructure used to connect process data, material response, and industrial
decision-making.
However, current ML applications in powder manufacturing still face several critical
barriers that limit their industrial reliability. Many datasets remain small, fragmented,
and biased toward simple geometries or laboratory-scale conditions. Labels such as
porosity, fatigue life, microstructure, and defect type often require destructive
or high-cost characterization. In-situ sensor data may be abundant, but their practical
value is limited by the quality of spatial and temporal registration with post-build
inspection results[18]. Powder reuse adds another layer of complexity because the same nominal process parameters
may lead to different outcomes when powder history, refresh ratio, or reuse strategy
changes[10,
11,
17]. As a result, a model may perform exceptionally well within one specific dataset
but fail completely when transferred to another powder lot, machine, alloy composition,
geometry, or post-processing route. This is fundamentally a data-organization problem
rather than a modeling limitation alone.
Several review articles have already discussed ML in AM and powder-based manufacturing
from important perspectives, including model families, process optimization, quality
control, in-situ monitoring, defect detection, property prediction, neural-network
methods, physics-informed learning, and closed-loop control[12–
16,
18,
19]. These reviews provide comprehensive summaries of how different algorithms have been
applied and what levels of predictive performance have been reported. However, much
of the existing literature is organized primarily around algorithms, AM processes,
or target prediction tasks. Less emphasis has been placed on powder-based manufacturing
as a lifecycle data-system problem, in which the powder state, process parameters,
monitoring signals, post-processing history, and final properties must remain traceable
within a single unified data thread. This distinction is vital because the central
challenge in powder manufacturing is not simply to predict a single property from
a table of parameters, but to preserve the complete material and process history required
to interpret, validate, and transfer that prediction.
The central thesis of this review is therefore that ML for powder-based manufacturing
should be evaluated not only by predictive accuracy, but also by its ability to maintain
lifecycle traceability. Under this framework, powder characterization, process monitoring,
post-processing records, and final testing are not viewed as separate data sources;
instead, they represent different layers of the same continuous manufacturing history.
A lifecycle data-system perspective makes it possible to ask more rigorous questions:
which powder-state variables are known, which process signals are measured, which
labels are truly reliable, which metadata are missing, and whether the resulting model
can be confidently transferred beyond the specific dataset used for training. This
framing also clarifies why uncertainty quantification, interpretability, domain shift,
and deployment feasibility are core requirements for reliable industrial deployment
rather than secondary implementation issues[10,
11,
20–
22].
Accordingly, this review makes four specific contributions. First, it reorganizes
ML applications in powder-based manufacturing according to the powder lifecycle, connecting
powder preparation, shaping or densification, in-situ monitoring, post-processing,
and final validation. Second, it classifies ML applications by task type—classification,
regression, sequence or spatiotemporal modeling, and optimization or inverse design—while
explicitly linking each task to the specific data sources it utilizes. Third, it analyzes
the primary barriers that prevent ML models from becoming reliable industrial tools,
including data scarcity, costly labeling, weak cross-stage traceability, domain shift,
poor generalization, and the trade-off between predictive accuracy and interpretability.
Finally, it systematically compares emerging strategies for overcoming these barriers,
including transfer learning, multi-fidelity modeling, physics-informed learning, and
hybrid learning approaches[13,
16,
21,
23–
25]. The goal is not to identify one universally superior algorithm, but to clarify which
modeling strategy is appropriate for a given powder dataset, lifecycle stage, and
decision-making objective.
The remainder of this review is organized as follows. Section 2 introduces the powder-to-part
workflow and summarizes representative ML applications across powder characterization,
process monitoring, defect detection, property prediction, post-processing-aware modeling,
and optimization. Figure 1 provides the lifecycle framing used in this review by linking the powder state, processing
history, monitoring signals, post-processing records, and final validation labels
within a single, unbroken data chain. Section 3 discusses the main barriers that limit
reliable ML deployment, including data scarcity, label fidelity, lifecycle incompleteness,
domain shift, and physical inconsistency. It then compares three data-efficient strategy
families: transfer-based learning, multi-fidelity modeling, and physics-informed or
hybrid learning. Section 4 concludes with future directions for traceable, uncertainty-aware,
and qualification-relevant ML implementations in powder-based manufacturing.
Fig. 1. Integrated ML workflow across the powder-manufacturing lifecycle, from raw
lifecycle data and engineered inputs to model selection, prediction, validation, and
industrial decision support.
2. CURRENT STATUS OF MACHINE-LEARNING APPLICATIONS IN POWDER-MANUFACTURING PROCESSES
2.1. The Powder Lifecycle as a Traceable Data System
In powder-based manufacturing, powder is better regarded as an evolving material state
than as a passive starting input. Its condition changes through production, sampling,
storage, handling, spreading or compaction, thermal exposure, reuse, reconditioning,
post-processing, and final inspection. As a result, the same nominal alloy, and even
the same nominal powder specification, may behave differently across batches when
particle-size distribution, morphology, surface chemistry, moisture level, contamination,
flowability, oxygen content, or reuse history changes. These variations affect powder
packing, energy absorption, melt-pool stability, densification, defect formation,
and the final microstructure and properties[4,
6,
7,
22]. This lifecycle view is important because powder-manufacturing data are generated
in separate but connected stages. Before processing, powder characterization provides
descriptors such as particle size, morphology, chemistry, apparent density, tap density,
flowability, and spreadability. During processing, powder is converted into an intermediate
state, such as a powder bed, green compact, melt pool, or sintering body. After consolidation
and post-processing, the same material history is evaluated through density, porosity,
microstructure, surface roughness, hardness, tensile properties, fatigue response,
or wear behavior. These data layers differ in format, scale, cost, and uncertainty,
but they describe the same powder-to-part route. If the data layers are not connected,
a measured property change cannot be assigned confidently to the feedstock, processing
route, machine condition, post-processing step, or their combined effect[6,
7,
22].
Powder characterization is therefore more than feedstock acceptance. Particle-level
descriptors such as size, shape, surface texture, internal porosity, and chemistry
must be considered together with bulk-level behavior such as flowability, apparent
density, tap density, rheology, and spreadability[4,
6]. Many standard powder tests remain useful, but they do not always reproduce the stress
state, spreading mechanism, environmental exposure, or machine-specific conditions
experienced during PBF. This is why a powder may pass common specification checks
and still produce unstable layers under a given recoating condition. Powder-evaluation
methods are most useful when selected in relation to the powder lifecycle, alloy system,
machine configuration, and intended process route6]. The same logic also applies to pressed powder metallurgy routes. A powder that shows
acceptable flowability or apparent density may still produce different green density,
density gradients, crack susceptibility, or sintering shrinkage when compaction pressure,
die geometry, lubrication, or particle morphology changes. For ML, green-density measurements,
compaction signals, and sintered-density labels should therefore be treated as linked
processing data rather than isolated tabular values[26–
29].
The powder-spreading step illustrates this point clearly. In PBF, the powder layer
is not simply a geometric input of prescribed thickness. It is a process state shaped
by powder cohesion, particle-size distribution, particle morphology, recoater type,
recoating speed, layer thickness, dosing condition, and machine-specific mechanisms[6,
7]. Layer quality can be quantified through descriptors such as layer-thickness deviation,
surface coverage, surface roughness, and packing density. For example, recent discrete
element method (DEM)-based work suggests that skewness and kurtosis of the layer surface
profile can serve as useful indicators of layer-quality variation[30]. For ML, these metrics matter because powder-bed images or recoating metrics should
not be treated as generic image data. They are process signatures that sit between
powder descriptors and final defects.
Powder reuse further strengthens the need for lifecycle tracking. Reused powder carries
a processing history that may include thermal exposure, spatter contamination, oxidation,
sieving, blending, storage, and repeated contact with the machine environment[10,
11]. These processes can alter particle morphology, surface chemistry, particle-size
distribution, flow behavior, packing response, and oxygen or nitrogen content. Different
reuse strategies also create different levels of history reconstruction. Single-batch
and collective-ageing approaches preserve clearer powder provenance, whereas top-up
and refreshing methods may preserve usability while making the exact exposure history
more difficult to reconstruct[10]. As a representative example, recent work on reused AlSi10Mg shows that recycling
count can interact with process parameters such as laser power and deposition thickness,
affecting void nucleation, indentation modulus, and wear behavior[17]. Reuse state needs to be recorded as a material-history variable, not treated as
background information.
Traceability remains equally important after consolidation. Final labels used for
ML, including density, porosity, microstructure, surface roughness, hardness, tensile
strength, fatigue life, and wear rate, are not direct outputs of a single process
parameter. They are outcomes of the full powder-to-part route. Heat treatment, HIP,
debinding, sintering, stress relief, machining, and surface finishing can all change
the relationship between the as-processed state and the final measured response[22]. Moreover, not all labels have the same fidelity. In laser powder bed fusion (LPBF)
density assessment, Archimedes measurements are relatively accessible and scalable,
but they provide bulk estimates and can be affected by sample geometry or measurement
conditions. Micrographic or computed tomography (CT)-based observations provide more
spatially resolved information on porosity and defects, but they are more labor-intensive
and harder to collect at scale[23]. This label hierarchy should be made explicit when training, comparing, or interpreting
ML models. The main bottleneck is therefore not only that powder-manufacturing datasets
are small. The more difficult problem is that powder qualification reports, machine
logs, in-situ monitoring files, post-processing records, and final testing results
are often stored as separate records. Multimodal sensor fusion is useful only when
different signals are linked to the same physical location, layer, specimen, or build
history. A representative powder-based AM example is the multimodal sensor-fusion
study of Petrich et al., where optical images, acoustic emission, multispectral emission,
scan-vector information, and machine logs were linked to CT-based flaw labels at the
voxel level[20]. In that workflow, the value of sensor fusion came from registration between sensor
features and post-build defect labels, not simply from using many sensors. The same
traceability logic applies to pressed powder metallurgy, where press-force or displacement
signals are useful for crack or compact-quality assessment only when the signal history
is connected to compact identity, tooling condition, sintering route, and final inspection
result.
A lifecycle dataset should preserve both material provenance and process provenance.
At minimum, this includes identifiers for powder origin and condition, machine and
sensor configuration, process parameters, post-processing history, and final characterization[6]. This does not mean that every study must measure every variable. Rather, unmeasured
variables should be recognized as missing parts of the manufacturing history. Missing
metadata are not neutral omissions. Missing powder, process, sensor, or post-processing
records define the boundary within which an ML model can be interpreted, validated,
or transferred[6,
11,
20]. The lifecycle view provides the basis for the ML discussion in the following section.
Classification, regression, spatiotemporal modeling, and optimization do not operate
on isolated datasets; they use different portions of the same lifecycle record. When
this record is incomplete, model accuracy may reflect dataset-specific correlations
rather than transferable process–structure–property relationships. When the record
is traceable, ML can be used more defensibly to screen powder quality, identify process
windows, connect in-situ signatures with defects, and support decisions about reuse
or post-processing. The goal is therefore not only to apply more complex algorithms,
but to build datasets in which the physical meaning of each input and label is retained.
2.2. ML Task Types Across the Powder-Manufacturing Data Thread
The lifecycle view in Section 2.1 can be translated into four practical ML task types:
classification, regression, sequence or spatiotemporal modeling, and optimization
or inverse design. These task types appear across powder metallurgy, binder jetting,
powder bed fusion, directed energy deposition, and related powder-processing routes.
The task categories are not restricted to additive manufacturing. Conventional powder
metallurgy also uses classification for quality screening, regression for density
or sintering prediction, time-dependent signals for compaction or process monitoring,
and optimization for composition, compaction, sintering, or post-processing design.
The common feature is that each task extracts a different type of decision from powder
state, processing conditions, monitoring signals, post-processing records, and final
quality measurements[12–
16,
27–
29]. Figure 2 organizes these task types within the powder-manufacturing data thread. Powder descriptors,
process variables, images, sensor signals, simulation outputs, and property measurements
can support different ML objectives, but the usefulness of each objective depends
on whether the input data and output labels remain traceable to the same powder batch,
specimen, build, layer, compact, or processing history.
Fig. 2. Major ML task categories in powder-manufacturing applications
For the powder-manufacturing data thread to support ML, powder descriptors, process
parameters, layer-level or compact-level monitoring signals, and ex-situ measurements
must be connected through common identifiers. These identifiers may include powder
lot, reuse cycle, build number, layer number, compact number, specimen location, tooling
condition, heat-treatment route, or measurement position. Without such links, traceability
remains a documentation concept rather than a usable structure for ML. This issue
is important in powder-based manufacturing because powder lot, reuse history, recoating
condition, compaction state, machine configuration, and post-processing route can
all influence the same final quality label[4,
6,
11,
15,
20,
21].
One important use of ML is screening and state recognition. In powder-processing applications,
this task is usually formulated as classification, where the output is a discrete
label rather than a continuous value. Classification includes anomaly detection, defect
detection, powder-layer assessment, powder quality screening, compact-quality classification,
swelling or shrinkage classification, crack detection, and pass/fail evaluation. The
input data may be powder-bed images, optical or thermal signals, acoustic signals,
photodiode traces, press-force signals, or tabular powder descriptors. The output
may indicate a nominal or abnormal layer, defective or non-defective region, cracked
or non-cracked compact, swelling or shrinkage behavior, or acceptable or unacceptable
powder condition[5,
20,
27,
31–
33].
The main limitation of classification is label reliability. In powder-based manufacturing,
a defect label is rarely caused by a single factor. A region marked as defective may
result from powder spreading, local energy input, powder reuse, contamination, atmosphere,
tooling condition, geometry, or post-processing. Rare defect occurrence, class imbalance,
subjective annotation, and inconsistent inspection thresholds can make a classifier
appear accurate within one dataset while failing on another powder lot, build, compact,
or machine. Classification outputs should therefore be treated as process-state indicators
unless the class labels are linked to objective validation such as CT, metallography,
dimensional inspection, mechanical testing, or other final quality measurements[20,
31–
33]. A warning label becomes more useful when the warning can be traced back to the powder
and process conditions that produced it.
A second group of ML applications aims to predict continuous responses. These applications
are usually formulated as regression problems. Typical regression targets include
density, porosity, melt-pool width or depth, surface roughness, hardness, tensile
strength, fatigue response, wear rate, sintering shrinkage, green density, and other
property-related quantities[26,
27,
29,
34,
35]. In PM and powder-based AM, regression inputs may include powder composition, particle-size
distribution, mass fraction, morphology, green density, compaction pressure, laser
power, scan speed, hatch spacing, layer thickness, sintering temperature, heat-treatment
condition, or HIP condition. Regression is attractive because many powder-processing
studies generate tabular datasets with measurable process variables and numerical
quality responses. The same tabular structure can also make regression misleading.
Many powder-processing datasets are small, highly correlated, and sparse in the true
process space. For example, laser power and scan speed are often varied together;
powder morphology and flowability are coupled; reuse history can change both chemistry
and spreading behavior; and post-processing can mask or amplify differences formed
earlier in the process. A regression model trained under these conditions can report
low error while learning experimental bias or dataset-specific correlations[13,
15,
22,
34,
35]. Regression predictions should therefore be checked not only by statistical metrics,
but also by physical consistency within the process–structure–property relationship.
When regression is used for transfer, extrapolation, or process recommendation, the
predicted response should be compared with micrographs, CT observations, powder-bed
data, compact-quality measurements, or known processing trends[22,
27,
29,
34–
36].
Forward and inverse modeling provide a useful distinction within prediction- and optimization-oriented
ML. Forward models predict process signatures, microstructure, build quality, properties,
or performance from powder and processing inputs. Inverse models start from a target
response and search for powder conditions or process parameters that may achieve it[13]. The inverse task is attractive because practical process development often requires
a window of acceptable conditions rather than a single prediction. However, inverse
outputs should be interpreted with care. If the training data come from one alloy,
powder lot, machine, geometry, or post-processing condition, the recommended conditions
may not remain valid after any of these factors changes[13].
A third class of ML applications deals with process evolution. Sequence and spatiotemporal
modeling become important when the data describe how a process changes over time or
across location rather than a single static state. Acoustic signals, emission spectroscopy,
photodiode traces, melt-pool videos, thermal histories, layerwise images, scan-vector
trajectories, press-force histories, and displacement signals can contain time- or
location-dependent information about powder spreading, compaction, melt-pool instability,
plume fluctuation, recoating disturbance, crack formation, and process drift[20,
31,
32,
37,
38]. These data are valuable because defects often develop through a sequence of local
events rather than appearing instantaneously in the final part or compact. A temporal
or spatiotemporal model can detect signatures of instability before the defect becomes
visible as porosity, distortion, cracking, or mechanical degradation. The difficulty
is data alignment. Sensors measure different physical phenomena at different sampling
rates, resolutions, and fields of view. A signal anomaly does not automatically correspond
to a defect unless the signal is registered to the correct layer, coordinate, scan
path, compact, specimen, and post-build inspection result. Multimodal monitoring should
therefore be treated as a data-alignment problem, not only as a model-architecture
problem[17,
28,
33,
34]. Time shifts, delayed responses, missing frames, sensor drift, and rare abnormal
events all weaken model transferability. In practice, the bottleneck is often not
the amount of sensor data, but the reliability with which sensor data are connected
to ground truth from CT, metallography, dimensional inspection, crack inspection,
or mechanical testing[20,
31,
36,
38].
Optimization-oriented ML extends prediction into decision support. Instead of asking
what output will result from known inputs, optimization searches for powder-production
settings, process parameters, powder conditions, alloy compositions, or post-processing
routes that are likely to meet target properties under constraints[13,
39–
41]. In this review, inverse design refers to decision-oriented search over feasible
powder, process, or post-processing conditions, rather than only the mathematical
inversion of a predictive model. Bayesian optimization is particularly attractive
because it can guide expensive experiments by using a surrogate model and uncertainty
estimate to decide which condition should be tested next[2,
41,
42]. Recent examples include multi-objective Bayesian optimization of LPBF Ti6Al4V processing
conditions and multi-response optimization of powder-based DED parameters, both showing
that process optimization in metal AM must balance density, surface quality, defect
formation, melt-pool geometry, and mechanical response rather than optimize a single
metric alone[40,
43]. This is valuable in powder-based manufacturing, where process-window development,
powder reuse qualification, and composition screening often require costly builds
and destructive characterization.
Most optimization workflows depend on a surrogate model, and their recommendations
are only as reliable as the data, assumptions, and constraints behind that surrogate.
The search space, feasible processing limits, powder-state variables, measurement
fidelity, and uncertainty estimates must be stated explicitly[13,
15,
21,
42]. Optimization recommendations are most defensible when they guide the next round
of experiments, narrow the parameter window, or rank candidate conditions. The same
recommendations are less defensible when presented as final process recipes without
experimental confirmation. For this reason, inverse design should be viewed as constrained
decision support under uncertainty rather than as a black-box replacement for process
qualification[26,
39]. Data-efficient strategies extend these four task types rather than replacing them.
Transfer learning can reduce target-domain labeling when related materials, machines,
sensors, or simulations provide useful source information. Multi-fidelity modeling
can combine lower-cost measurements or simulations with fewer high-fidelity labels.
Physics-informed and hybrid learning can use physical priors, simulations, or mechanistic
descriptors to improve plausibility under sparse-data conditions. These three strategy
families are discussed in Section 3 because their usefulness depends on source–target
similarity, fidelity hierarchy, physical validity, and lifecycle traceability[15,
16,
23–
25,
42,
44–
46].
The four task types should be viewed as complementary operations on the same powder-manufacturing
data system. Classification is strongest for rapid screening and state recognition.
Regression is strongest for continuous property or process-response prediction. Sequence
and spatiotemporal models are strongest for tracking process evolution. Optimization
and inverse design are strongest for constrained decision-making and experimental
planning. In practice, the task types also interact: classification outputs may become
inputs to regression, sequence features may support defect prediction, and optimization
usually depends on a surrogate predictive model. From a materials-science perspective,
the central issue is not only predictive accuracy. The central issue is whether powder
history, process evolution, and final validation remain connected through a traceable,
physically interpretable, and transferable data chain.
2.3. Representative ML Applications Across the Powder-Processing Workflow
The task categories discussed above become more meaningful when they are placed back
into the manufacturing workflow. In powder-based manufacturing, ML is not applied
at a single point. It appears at the feedstock stage, during layer formation or densification,
in process monitoring, after post-processing, and finally in property evaluation or
process optimization. These applications differ in input format and modeling objective,
but they share a common requirement: predictions must remain traceable to the powder
and process history from which the data were generated.
At the feedstock stage, ML can be used to model and optimize powder production itself.
Powder yield, particle-size distribution, sphericity, satellite formation, and morphology
are not only powder-supplier concerns; they define the starting condition for spreading,
packing, melting, sintering, and final property development[4,
6]. Tamura et al. studied gas-atomized Ni–Co-based superalloy powders for turbine-disk
applications and used a Gaussian-process surrogate model with Bayesian optimization
to select melt temperature and gas pressure[41]. Starting from three initial experiments and three optimization cycles, the study
reported a qualified powder fraction of 77.85% for particles smaller than 53 μm, together
with an estimated production-cost reduction of about 72% relative to commercial powder[41]. The modeled target was not a final part property, but the powder state itself. This
places ML upstream in the powder lifecycle, before spreading, melting, sintering,
or post-processing begins. A different set of applications begins once the powder
is spread, compacted, or otherwise converted into a process state. Powder delivery
is also an important process-state variable in powder-fed directed energy deposition.
Jung et al. showed that variations in powder line density altered bead geometry, grain
morphology, anisotropy, and mechanical properties in L-DED-fabricated STS316L, even
under comparable energy-density conditions[47]. In laser powder bed fusion (LPBF), the powder layer is often treated as if it were
defined only by nominal layer thickness, but actual layer quality depends on particle
cohesion, size distribution, recoater motion, surface coverage, surface roughness,
and packing density[6,
7]. Discrete element method (DEM)-based studies have shown that layer quality can be
described using measurable descriptors such as layer-thickness deviation, surface
coverage ratio, root-mean-square roughness, packing density, skewness, and kurtosis[30]. These descriptors turn the powder bed from a visual observation into a measurable
process signature. Figure 3 illustrates how layer-wise or powder-bed monitoring data are converted into model
inputs and linked to process-state or defect labels. This connection is central to
evaluating whether an ML model detects a physically meaningful anomaly or only separates
patterns within a specific dataset.
Fig. 3. Example of ML-based in-situ monitoring and anomaly-pattern detection during
LPBF [5].
Computer-vision methods have been used to extract such signatures directly from powder-bed
images. Scime and Beuth proposed an early LPBF monitoring workflow in which post-recoating
images were processed to detect and classify powder-bed anomalies[5]. Their training database contained 2402 labeled image patches covering anomaly-free
regions and six anomaly classes[5]. In that workflow, a powder-bed image was first converted into numerical visual descriptors.
Handcrafted visual features are manually designed descriptors, such as intensity,
texture, edge, or shape-related information, that represent the appearance of a local
image region. A bag-of-visual-words representation treats recurring local visual patterns
as “visual words” and summarizes their occurrence in an image as a compact feature
vector. Scime and Beuth used k-means clustering as an unsupervised method to group
similar local feature descriptors and define the visual-word vocabulary used for anomaly
classification[5]. The resulting feature vectors were then used for powder-bed anomaly classification.
This workflow is useful because it shows how recoating defects can be converted from
visual observations into trainable layerwise features, rather than being treated only
as qualitative images. Later image-based work increased both image resolution and
annotation scale. Fischer et al. used more than 45,000 annotated powder-bed anomalies
and reported 99.15% classification accuracy with class F1-scores between 97.85% and
99.71% under the best imaging and model conditions[33]. These results support the use of powder-bed imaging for layer-quality monitoring,
but they also define the boundary of the evidence. Performance remains sensitive to
image resolution, lighting, anomaly definition, powder reflectivity, and the representativeness
of expert labels.
More recent monitoring studies have moved from single-image analysis toward multimodal
sensor fusion. Instead of relying on one signal, these workflows combine several records
from the same build. Layerwise optical images refer to images captured at individual
build layers, typically after powder recoating or after laser exposure, so that local
layer disturbances can be linked to a specific layer and build location. These images
can be combined with acoustic signals, multispectral emissions, scan-vector information,
and machine logs, and then correlated with post-build inspection results such as CT-detected
flaws, metallographic porosity, or location-specific quality measurements[20,
31,
32]. Petrich et al. provide a useful scale reference for this type of work: their multimodal
sensor-fusion model used 168,574 voxel-level samples, four-fold cross-validation,
and reported 98.5% binary flaw/no-flaw classification accuracy against CT-linked labels[20]. In that workflow, the reported accuracy was meaningful because sensor features were
registered to build location, layer number, scan-vector information, and final flaw
labels. Without this registration, a sensor feature may still correlate with a defect,
but the physical meaning of the prediction becomes difficult to defend.
ML has also been used to predict part quality and mechanical properties from process
and material variables. In these cases, the model usually receives structured inputs
such as powder composition, particle size, energy input, scan speed, hatch spacing,
layer thickness, build orientation, heat-treatment condition, or HIP parameters, and
then predicts density, porosity, hardness, tensile strength, elongation, fatigue life,
or wear behavior[13–
15,
27,
34,
35]. Destructive characterization is expensive, so validated property-prediction models
can reduce the number of experiments needed to screen a process window. Their reliability,
however, depends on whether powder lot, reuse state, build location, specimen location,
and post-processing condition are recorded. A recent LPBF study on 3.3% Si electrical
steel also combined ML and explainable AI to model density, surface roughness, and
hardness as functions of laser power, scan speed, and scanning angle, showing that
ML-based process optimization can be linked directly to powder-bed-fusion process
variables and experimentally measured quality indicators[48].
Post-processing illustrates this issue particularly well. Final performance is often
not determined by the build step alone. Heat treatment, stress relief, and HIP can
alter residual stress, microstructure, porosity, surface condition, and mechanical
scatter[18,
19,
22,
34]. Yang et al. developed an artificial neural network model for LPBF-fabricated Ti–6Al–4V
components in which as-printed properties and HIP parameters were included as inputs
for predicting final tensile properties[34]. The model predicted yield strength and ultimate tensile strength more reliably than
elongation: 87.5% of yield-strength predictions and 100% of ultimate-tensile-strength
predictions were within 5% error, whereas only 62.1% of elongation predictions were
within 10% error[34]. This uneven performance is important for lifecycle modeling. Strength can often
be captured by processing and heat-treatment variables, but elongation is more sensitive
to hidden defects, surface condition, residual porosity, and source-study variability.
HIP or heat treatment should therefore be modeled as a downstream transformation of
the as-built state, not left as background experimental information. Process-window
construction is another area where ML has become useful. In LPBF, density is often
used as a first indicator of process quality, but density measurements can differ
in cost, resolution, uncertainty, and spatial representativeness. Song et al. used
multi-fidelity Gaussian-process modeling to combine 60 low-fidelity Archimedes measurements
with 25 high-fidelity micrographic observations for constructing a process window
for a thin-walled LPBF structure[23]. Archimedes measurements provided lower-cost bulk-density coverage, whereas micrographic
observations supplied more spatially resolved porosity information. The workflow reflects
a practical measurement problem in powder processing: the most accessible data are
not always sufficient to resolve local defects, while the most spatially informative
labels are often too costly to collect at scale.
Powder reuse has recently become another application area where ML can support decision-making.
Reuse can change powder state through oxidation, contamination, agglomeration, changes
in particle-size distribution or morphology, and altered flow behavior[10,
11]. In an experimental–ML study on reused AlSi10Mg powder in LPBF, laser power, deposition
thickness, and reuse count were used to predict void nucleation, indentation modulus,
and wear behavior[17]. The study varied laser power at 280, 380, and 480 W, reuse cycles at 5, 7, and 9
cycles, and deposition thickness at 35, 60, and 85 μm; the reported optimum was 380
W, 35 μm, and 7 reuse cycles, with 1.94% void nucleation, a wear rate of 0.97 × 10-4 mm3/Nm, and an indentation modulus of 115.15 GPa[17]. The exact optimum should not be generalized beyond the tested material and machine
conditions. The reusable lesson is not the specific parameter set, but the treatment
of reuse count as an explicit process-history variable. A model that omits reuse history
may assign porosity or mechanical-property changes to laser parameters even when powder
aging or contamination contributes to the response.
Across these workflow stages, optimization and inverse design represent the decision-oriented
use of ML. Instead of only predicting the consequence of a given condition, these
approaches search for powder-production settings, process parameters, compositions,
or post-processing routes that are likely to meet a target response[13,
15,
21,
39–
41]. Bayesian optimization is particularly attractive when each experiment is expensive,
as in gas atomization, process-window development, and mechanical-property qualification[13,
42,
43]. Digital-twin-oriented frameworks extend this idea by using time-series models and
uncertainty-aware optimization to recommend process adjustments during manufacturing[21]. These outputs should be treated as recommendations for the next experimental step,
not as final recipes. Their reliability depends on the surrogate model, the search
space, the uncertainty estimate, and the lifecycle data used for training. Viewed
along the workflow, these examples show that ML applications in powder-based manufacturing
now cover a wider part of the powder-to-part chain. Powder-production optimization,
recoating-anomaly detection, multimodal monitoring, property prediction, post-processing
analysis, reuse assessment, and process-window construction represent different points
along the same data chain. When powder history, process signatures, post-processing
records, and final labels are connected, ML can support screening, prediction, and
decision-making with a clearer validation boundary. When those links are missing,
even a model with strong internal performance may remain difficult to interpret, reproduce,
or transfer.
Additional representative cases are summarized in Table 1. The table extends the workflow-based discussion by comparing studies according to
workflow stage, data and model type, scale and validation design, reported metric,
and main limitation. This structure is intended to show not only where ML has been
applied, but also how strongly each result is supported by data scale, label quality,
and validation boundary.
Table 1. Representative ML and data-driven studies across the powder-based manufacturing
workflow and data-efficient modeling strategies
|
Stage
|
Data type & model
|
Scale & Validation
|
Reported metric
|
Key limitation
|
References
|
|
Powder manufacturing
|
Gas-atomization parameters; GP-based Bayesian optimization
|
3 initial experiments + 3 optimization cycles; external atomizer validation NR
|
Qualified yield (<53 μm) = 77.85%; cost reduction ≈ 72%
|
Validated only for the tested Ni–Co alloy powder and gas-atomization setup; transfer
to other alloys or atomizers was not reported.
|
Tamura et al.[41]
|
|
Feedstock QC
|
VIS/NIR HSI; spectral dictionary; ML classification; band selection
|
5 original powders + 8 mixed samples; external validation NR
|
Contamination characterized down to 1% under surface /pixel-size conditions
|
Surface exposure and pixel mixing dependent
|
Yan et al.[38]
|
|
Powder spreading / layer quality
|
DEM + Taguchi DoE; powder-bed images; computer vision / pretrained deep-learning models
|
Avrampos et al.: DEM /Taguchi-based layer-quality analysis; Scime et al.: 2402 labeled
image patches
|
Avrampos et al.: optimum deviation −12.9% / −3.5%; Scime et al.: anomaly-free regions
and six anomaly classes classified from post-recoating images
|
Lighting, resolution, and anomaly-definition dependent
|
Avrampos et al.[30]; Scime et al.[5]
|
|
In-situ flaw / process-state monitoring
|
Multimodal sensors; thermography; photodiode signals; CNN; CT/XCT labels
|
Petrich: n = 168,574 voxels, 4-fold CV; Cao: 36 tracks, 80:20 + 10-fold CV
|
Petrich: Acc = 98.5%; Oster: Acc = 0.96, F1 = 0.86; Cao: Acc = 95.81%, 15 ms/sample
|
Cross-machine/material /powder-lot transfer limited or NR
|
Petrich et al.[20]; Oster et al.[31]; Cao et al.[32]
|
|
PM density / sintering prediction
|
Materials descriptors; composition/powder/process variables; regression models
|
Zhang: 223 HVC entries + 9 validation instances; Kamal: n = 460; Asnaashari: n = 210,
80:20 split
|
Zhang: error <2%; Kamal: RF MAE = 0.024, validation MAE = 1.82%; Asnaashari: R = 0.989,
RMSE = 0.016
|
Alloy-family and route specific
|
Zhang et al.[26]; Kamal et al.[27]; Asnaashari et al.[29]
|
|
PM classification / crack quality
|
Process descriptors; hydraulic-press sensor features; RF / ensemble classifiers
|
Kamal: n = 211, 70:30 split + 5-fold CV; Mustafa: production press-signal data
|
Kamal: RF Acc = 0.92; Mustafa: best Acc up to 99%
|
Geometry, tooling, and crack-location dependent
|
Kamal et al.[28]; Mustafa et al.[36]
|
|
Composition-based printability
|
Composition, elemental, and process descriptors; RF / GB / NN
|
Balling n = 267; porosity n = 138; external validation NR
|
Balling NN Acc = 92.3%; porosity RF R2 = 0.971, RMSE = 0.109
|
Powder state and machine history incomplete
|
Roy et al.[49]
|
|
Powder reuse / properties
|
Reused AlSi10Mg; LPBF parameters; ridge regression / RF
|
280/380/480 W; 5/7/9 reuse cycles; 35/60/85 μm; external validation NR
|
Optimum: 380 W, 35 μm, 7 cycles; void = 1.94%; wear = 0.97 × 10-4 mm3/Nm; modulus = 115.15 GPa
|
Powder-lot and reuse-protocol transfer NR
|
Murugesan et al.[17]
|
|
Post-processing-aware properties
|
Literature-derived LPBF Ti6Al4V + HIP database; ANN
|
Validation by prediction-error ranges; prospective validation NR
|
YS: 87.5% within 5% error; UTS: 100% within 5%; elongation: 62.1% within 10%
|
Elongation sensitive to hidden printing defects
|
Yang et al.[34]
|
|
Multi-fidelity process window
|
LF Archimedes density + HF micrography; MF-GPR
|
60 LF + 25 HF; LOOCV
|
LOOCV error: LF 7.88 / HF 0.88 → MF 0.19; cost ≈ USD 2180 vs USD 6800 all-HF
|
Thin-wall LPBF density/process-window specific
|
Song et al.[23]
|
|
Transfer learning / model reuse
|
Active cross-platform transfer; HTC-to-HTE transfer learning
|
Zheng: 265 M290 + 36 AM250 + 32 DMP350; AM250 transfer case with active/random/from-scratch
comparison. Li: 302 HTE data; extrapolation validation; 105 screened compositions.
|
Zheng: target RMSE ≈ 35 MPa in the AM250 case; Li: extrapolation Acc = 90.48%, Recall
= 95.06%; γ′ MAPE = 2.21%, 4.28%, 5.13%.
|
Source–target similarity dependent
|
Zheng et al.[42]; Li et al.[50]
|
|
Physics-informed learning
|
PINN; architecture-driven physics-informed learning
|
Tiwari: n = 347, test n = 52; Ghungrad: n = 1000, 80:20 split
|
Tiwari: MAPE = 3.8%, 4.7%, 3.1%, 1.9%; Ghungrad: MAPE = 2.85%, R2 = 0.936
|
Limited to regimes where thermal assumptions hold
|
Tiwari et al.[24]; Ghungrad et al.[45]
|
Abbreviations: PM, powder metallurgy; ML, machine learning; LPBF, laser powder bed
fusion; GP, Gaussian process; MF-GPR, multi-fidelity Gaussian process regression;
CNN, convolutional neural network; RF, random forest; GB, gradient boosting; NN, neural
network; ANN, artificial neural network; DEM, discrete element method; DoE, design
of experiments; CT, computed tomography; XCT, X-ray CT; HSI, hyperspectral imaging;
VIS/NIR, visible/near-infrared; LF/HF, low fidelity/high fidelity; CV, cross-validation;
LOOCV, leave-one-out cross-validation; HVC, high-velocity compaction; HTC/HTE, high-throughput
calculation/high-throughput experiment; HIP, hot isostatic pressing; PINN, physics-informed
neural network; Acc, accuracy; F1, F1-score; MAE, mean absolute error; MAPE, mean
absolute percentage error; RMSE, root mean square error; R, correlation coefficient;
R2, coefficient of determination; YS, yield strength; UTS, ultimate tensile strength;
NR, not reported; γ′, gamma-prime phase.
3. ISSUES AND KEY CHALLENGES IN APPLYING ML TO POWDER-BASED MANUFACTURING
3.1. Data Scarcity, Label Fidelity, and Lifecycle Incompleteness
Data scarcity remains a major barrier to reliable ML in powder-based manufacturing.
The issue is not only the number of experiments. Available datasets are often sparse,
uneven, weakly labeled, and split across different stages of the powder lifecycle.
A dataset may contain process parameters without powder history, in-situ monitoring
signals without post-build validation, or final mechanical properties without sufficient
information about post-processing and specimen location. Under these conditions, a
model may learn a local statistical association without capturing the physical route
through which powder state, process evolution, and final properties are connected[12,
13,
15,
16].
The constraint is built into the experimental route. A single labeled data point may
require powder preparation, powder characterization, spreading or compaction, melting
or sintering, heat treatment, machining, microscopy, CT, or destructive mechanical
testing. For high-value alloys, the cost is increased by expensive powder batches,
limited machine access, and the difficulty of reproducing failed builds under the
same conditions. Many ML studies therefore rely on narrow process windows, simplified
geometries, small parameter sets, or data collected on one material and one machine[13,
32]. Such datasets are useful for feasibility studies, but they rarely cover the variability
expected in production.
The same limitation appears outside additive manufacturing. In conventional powder
metallurgy, sintering outcomes depend on powder characteristics, alloy chemistry,
green density, compaction pressure, sintering atmosphere, heating rate, holding time,
and sintering temperature. Recent ML studies on sintered bronze and Cu-based powder
metallurgy alloys used datasets with a few hundred samples, including 460 data points
for bronze/Cu-based density prediction, 211 samples for Cu–Sn swelling or shrinkage
classification, and 210 data points for Cu–Al sintered-density prediction[27–
29]. These studies show that useful models can be built from curated experimental and
literature-derived datasets. Their transferability, however, remains bounded by the
alloy families, powder descriptors, green-density range, and thermal histories represented
in the training data. Image-based monitoring shows another form of the data problem:
a larger image dataset does not automatically remove label uncertainty. Scime and
Beuth used post-recoating LPBF images to detect and classify powder-bed anomalies
such as recoater streaking, debris, and part-related failures[5]. Their database contained 2402 labeled image patches, and the workflow converted
powder-bed images into handcrafted visual features and bag-of-visual-words representations
for anomaly classification[5]. Later image-based monitoring work increased the annotation scale to more than 45,000
powder-bed anomalies and reported high classification performance under controlled
imaging conditions[33]. This progression shows that larger labeled image datasets can improve training stability,
but the evidence still depends on image resolution, illumination, powder reflectivity,
recoater configuration, anomaly definition, and expert-label consistency. A model
trained under one imaging setup may therefore require adaptation before being used
with another machine, alloy system, powder lot, or recoating condition[15,
16,
20,
33].
Label scarcity is not limited to images. In-situ monitoring can generate large volumes
of acoustic, optical, spectral, thermal, or photodiode data, but supervised learning
requires these signals to be linked to trustworthy ground truth. A signal cluster
or latent representation should not be interpreted as lack of fusion, keyholing, cracking,
or contamination unless the signal is supported by independent evidence from CT, metallography,
mechanical testing, or another physically meaningful label[20,
31,
32,
37]. Rare events such as lack-of-fusion pores, keyhole defects, cracks, severe recoating
streaks, contamination events, and abnormal powder-layer disturbances may be critical
for qualification, but they occur much less frequently than nominal regions. A classifier
can therefore show high overall accuracy while missing the events that matter most
for safety or reliability. Accuracy should be reported together with class-wise recall,
precision, confusion matrices, or defect-specific performance, especially when the
positive class represents a rare but critical defect[5,
13,
20,
31,
32].
A related coverage problem appears in composition-based printability datasets. Roy
et al. compiled data for balling and porosity prediction from alloy composition and
process descriptors, with 267 data points for balling and 138 data points for porosity[49]. Such datasets are useful for rapid screening across composition and process space,
but their reliability depends on whether rare alloy families, defect modes, powder
states, and machine conditions are represented rather than merely interpolated from
nearby examples. Label fidelity also varies across measurement methods. Bulk density,
visual inspection, or a limited set of tensile tests can provide accessible scalar
labels, but they may not resolve local porosity, defect morphology, surface-connected
flaws, or microstructural variation. Song et al. addressed this problem in LPBF process-window
construction by combining broader-coverage Archimedes density measurements with more
labor-intensive micrographic observations through multi-fidelity Gaussian-process
modeling[23]. The example is useful here because it separates label availability from label fidelity.
Broader-coverage measurements help map the process space, while spatially resolved
observations are needed where local porosity or defect morphology controls the final
interpretation. Applicability domains should be treated as part of validation, not
as a separate post-analysis. A model trained within a curated dataset can report low
error while remaining valid only for a narrow range of compositions, particle sizes,
compaction pressures, sintering conditions, or heat-treatment histories. PM sintering
studies illustrate this point. Some studies include experimental validation under
selected alloy and processing conditions, whereas others use multiple error metrics
and leverage-based outlier analysis to identify data outside the model domain[27,
29]. Validation therefore cannot be reduced to a single accuracy or error value. It must
also specify where the input features, alloy systems, and processing conditions remain
physically meaningful. Random train–test splitting is often too weak for this setting.
Specimens from the same build, powder lot, geometry, imaging condition, or experimental
campaign can appear in both training and test sets, making the model look more general
than it is. Stronger tests include leave-one-build, leave-one-machine, leave-one-powder-lot,
grouped experiment splits, or external-dataset validation[15,
16,
20]. These validation designs are harder to satisfy, but they better reflect the way
ML models fail in powder-based manufacturing: not by small random errors within one
dataset, but by distribution shifts across powder lots, machines, sensors, geometries,
and post-processing routes.
History dependence makes the input–output relationship non-unique. The same final
density or strength may result from different combinations of powder state, thermal
exposure, post-processing, and defect morphology. Conversely, the same nominal process
parameters may produce different outcomes when the powder batch, reuse condition,
recoating response, atmosphere, or machine state changes[10,
20,
21,
46,
51]. Without lifecycle identifiers, apparent accuracy can appear high because the model
has learned batch-specific, geometry-specific, or machine-specific correlations rather
than a transferable process–structure–property relationship.
The data-scarcity problem therefore cannot be solved by increasing dataset size alone.
Additional data are useful only when they add diversity, improve label fidelity, preserve
provenance, and cover relevant powder and process states. Poorly aligned or weakly
documented data can increase sample count without improving model reliability. This
bottleneck motivates the strategies discussed in the following subsections. Transfer-based
methods reuse information from related materials, machines, sensors, or simulations
when the target domain has limited labels. Multi-fidelity methods combine lower-cost
or broader-coverage measurements with fewer high-fidelity or spatially resolved observations.
Physics-informed and hybrid methods constrain learning using prior knowledge from
heat transfer, fluid flow, densification, thermodynamics, or process simulation. These
strategies improve sample efficiency in different ways, but none removes the need
for traceable lifecycle data.
3.2. Domain Shift and Transfer-Based Generalization
The limitations described above lead to a second problem: a model trained on one powder-processing
domain rarely moves unchanged to another. Here, a domain refers to the material, powder
lot, machine, sensor configuration, geometry, process window, post-processing route,
and labeling method that define how the data were generated. Domain shift occurs when
one or more of these conditions change. Across powder routes, this shift is common
rather than exceptional. A model calibrated on one LPBF machine may not preserve its
error level on another machine, even for the same alloy, because gas flow, beam delivery,
recoating behavior, chamber geometry, and scan-control implementation are not identical[13,
15,
20,
42].
Random train–test splits can hide this problem. If samples from the same build campaign,
powder batch, or machine appear in both training and test sets, the reported error
may describe interpolation within one experimental context rather than generalization
to a new context. Domain-aware validation is stricter. It asks whether the model still
works when the powder lot, machine, alloy family, specimen geometry, sensor setup,
or post-processing route changes. This distinction is central for industrial use because
deployment almost always involves a target domain that is not identical to the training
domain. Domain-aware validation is therefore a diagnostic step before transfer learning
is attempted. It tests whether the model is still reliable after a change in powder
lot, machine, alloy family, specimen geometry, sensor setup, or post-processing route,
rather than only within a random split of the original dataset[13,
15,
42]. Transfer learning offers one route through this problem. In this review, transfer
learning refers to using information from a related source domain to improve modeling
in a target domain where labeled data are scarce. The transferred information may
be a pretrained representation, a source model, a calibrated simulation prior, or
a source dataset that is adjusted using a small number of target-domain measurements.
The method is attractive for powder-based manufacturing because target-domain labels
are often expensive: tensile tests, CT inspection, metallography, creep tests, and
high-temperature validation cannot be generated at the scale required by conventional
data-hungry models. Cross-platform LPBF property prediction provides a clear example.
Zheng et al. studied active transfer learning for ultimate tensile strength prediction
across three LPBF platforms: 265 data points from an EOS M290, 36 from a Renishaw
AM250, and 32 from a 3DSystems DMP350[42]. The model used laser power, scan speed, and hatch distance as inputs, and treated
cross-platform differences as prediction errors that could be learned from a small
set of target-platform tensile tests. In one AM250 transfer case, the model reached
a target RMSE level of approximately 35 MPa with fewer actively selected target samples
than random transfer or from-scratch modeling[42]. The gain is best interpreted as a reduction in target-machine labeling burden, not
as proof of universal cross-platform generalization. Each added target label was a
tensile-test result from the target platform, and the transferred model still depended
on how well the source and target platforms shared the same process–property relationship.
The same study also reported large transfer errors near regions associated with severe
process defects[42]. Transfer learning can therefore reduce the number of target-domain experiments when
the source and target platforms remain physically comparable. Transfer learning cannot
compensate for a source model when the target condition moves into a different regime,
such as lack-of-fusion, keyhole instability, or another defect-dominated response.
A second transfer setting appears in computation-to-experiment alloy design. In nickel-based
powder-metallurgy superalloy development, Li et al. combined high-throughput calculation,
diffusion-multiple experiments, and transfer learning to connect computational microstructure
information with sparse experimental data[50]. Their framework transferred from a high-throughput calculation domain to high-throughput
experimental data, and then used the calibrated microstructural predictions as part
of downstream property modeling and alloy screening. The reported extrapolation test
for topologically close-packed phase classification reached 90.48% accuracy and 95.06%
recall, and the γ′ feature regression models reported extrapolation MAPE values of
2.21%, 4.28%, and 5.13%[50]. The study also screened 105 candidate compositions after calibration with sparse experimental data[50]. In this case, transfer learning did not simply reuse a model; it used experimental
data to correct a broader computational prior. Monitoring models provide another transfer
setting. Image, acoustic, thermal, photodiode, or spectrogram-based representations
may retain low-level signal features across related materials or machines, but the
defect meaning of those features still depends on sensor configuration, process regime,
powder condition, and post-build validation. A feature that separates process states
in one dataset may not correspond to the same defect class after a change in powder
absorptivity, melt-pool stability, recoater behavior, sensor field of view, or labeling
procedure[20,
31,
32,
37].
These examples represent different forms of transfer. The LPBF case transfers across
machines that produce the same nominal material, while the PM superalloy case transfers
from computation-rich data to experiment-scarce alloy design. Monitoring transfer
often lies between these cases because the learned representation may be portable,
while the associated defect label may not be. Cross-platform LPBF transfer assumes
that source and target machines share enough process physics for calibrated error
correction to remain meaningful. Computation-to-experiment transfer assumes that the
computational source domain contains useful trends even when it is biased relative
to experimental observations. In each case, the target-domain calibration data define
the boundary of trust. Negative transfer remains the main risk. It occurs when information
from the source domain degrades target-domain prediction. In powder-based manufacturing,
negative transfer can arise from changes in powder morphology, oxygen content, reuse
history, recoater behavior, melt-pool regime, sintering route, or heat-treatment response.
A model transferred across materials may preserve image features while losing defect
meaning; a model transferred across machines may preserve parameter trends while shifting
absolute property values; a model transferred from simulation may inherit systematic
errors in thermodynamics, heat transfer, or densification kinetics. These failures
are not always visible from global accuracy alone. A useful transfer study should
therefore report more than a target-domain metric. It should specify the source domain,
target domain, number of target calibration samples, selection strategy for those
samples, validation split, and conditions where transfer fails. For powder-processing
applications, applicability-domain analysis is especially useful because the costliest
errors occur at the edge of the tested space: new powder lots, new machines, higher
reuse cycles, extreme energy densities, unfamiliar geometries, or post-processing
histories outside the training record[15,
16,
42,
50]. Transfer learning is best viewed as a way to reduce target-domain labeling burden,
not as a substitute for target-domain validation. It is most defensible when the source
and target domains share a physically plausible relationship and when the remaining
mismatch is measured rather than assumed away. When the difference between data sources
is better described primarily as a fidelity hierarchy rather than as a source–target
domain shift, multi-fidelity modeling may be the more appropriate framing because
it treats lower-cost and higher-fidelity evidence as related but explicitly different
evidence streams. That distinction motivates the next subsection[23,
42,
50].
3.3. Multi-Fidelity Learning for Data-Efficient Modeling
A fidelity hierarchy exists when several data sources describe a related response
but differ in cost, resolution, reliability, or physical completeness. Unlike domain
shift, which concerns transfer between different data-generating domains, multi-fidelity
learning focuses on how lower-cost evidence and higher-fidelity evidence can be combined
within a related modeling task. In powder-based manufacturing, this situation appears
frequently: bulk density can be measured faster than local porosity, simplified simulations
can cover a wider parameter space than experiments, and rapid screening tests can
rank candidate conditions before high-cost validation is performed[23,
52,
53].
Fidelity is target-dependent. A data source should not be called high fidelity in
an absolute sense; it is higher fidelity only relative to a specified target, measurement
scale, and decision. In this review, fidelity is also used in a measurement-hierarchy
sense, not only in a simulation-versus-experiment hierarchy. For example, Archimedes
density may be adequate for evaluating bulk densification, but it becomes lower-fidelity
information when the target is local pore morphology in a thin wall because it provides
low-cost bulk-density data without resolving local defect morphology. Micrographic
observation is treated as higher-fidelity information in that context because it provides
more spatially resolved defect information, although it requires greater effort and
cost[23]. More generally, multi-fidelity learning can combine lower-cost simulations, reduced-order
models, literature-derived data, rapid screening measurements, or lower-resolution
experiments with smaller amounts of higher-fidelity experimental evidence[23,
52,
53]. The LPBF process-window study by Song et al. gives a compact example. The study
used Archimedes density as the low-fidelity source and micrographic observation as
the high-fidelity source for thin-walled LPBF specimens[23]. The dataset contained 60 Archimedes measurements and 25 micrographic observations
for the thin-wall case[23]. Leave-one-out cross-validation showed the effect of combining the two sources: the
reported error was 7.88 for the low-fidelity model, 0.88 for the high-fidelity model,
and 0.19 for the multi-fidelity model. The reported measurement cost was also lower:
approximately USD 2180 for the mixed LF/HF workflow, compared with about USD 6800
if all measurements were collected at high fidelity[23].
The technical value of the method lies in how the two sources are linked. In multi-fidelity
Gaussian-process modeling, the lower-fidelity model supplies a broad trend across
the process window, while the model uses the higher-fidelity data to learn the discrepancy
between that trend and the higher-fidelity response. This structure is useful only
when the fidelity levels are correlated but not identical. If the lower-fidelity source
has no stable relationship with the high-fidelity target, additional low-fidelity
data can make the model more confident without making the prediction more reliable[23]. Analytical models, reduced-order thermal simulations, discrete element method simulations,
CALPHAD calculations, and finite-element analyses can explore wider process or composition
spaces than experiments. These models should not replace measurement; they are most
useful when they provide trends that can be calibrated against higher-fidelity or
application-relevant data[52,
53]. A coarse thermal model, for example, may be useful for ranking scan conditions,
but it cannot by itself validate local defect morphology or final tensile performance.
In the same way, a powder-spreading simulation may identify trends in layer uniformity
while still requiring experimental imaging, density measurement, or part-quality inspection
before it can support process qualification.
The failure mode is label misalignment. Bulk density, local porosity, CT-detected
flaw volume, micrographic area fraction, melt-pool features, and mechanical performance
are related, but they are not interchangeable labels. A model may combine them successfully
only if the spatial scale, measurement location, specimen identity, build history,
and post-processing route are traceable. Without those links, multi-fidelity learning
can mix responses from different physical levels of the workflow and produce a process
window that looks precise but is difficult to interpret. Multi-fidelity learning also
changes how experiments should be allocated. Instead of collecting every label at
the highest available fidelity, the model can use broad low-fidelity coverage to locate
regions where additional high-fidelity data are most valuable. This connects multi-fidelity
modeling with active learning and Bayesian experimental design. The main purpose is
to spend high-fidelity measurements where they reduce uncertainty in the response
that matters for the decision. High-throughput computation and experiments follow
the same logic when they are used carefully. Broad CALPHAD-based calculations, reduced-order
simulations, or high-throughput screening experiments can cover large regions of composition
or process space, while smaller experimental datasets provide correction and validation.
The nickel-based PM superalloy example discussed in Section 3.2 illustrates a neighboring
case: broad computational data supplied coverage, while diffusion-multiple experimental
data provided calibration for microstructural prediction[50]. The lesson for multi-fidelity modeling is similar: large imperfect datasets become
more useful when their bias is measured against smaller, more reliable observations.
For powder-based manufacturing, multi-fidelity learning is most useful when three
conditions are satisfied. First, the fidelity levels must refer to a common target
or to targets with a clearly defined mapping. Second, the low-fidelity source must
preserve a trend that is relevant to the high-fidelity response. Third, sample identity
and process history must remain traceable across fidelity levels. When these conditions
are missing, the model may gain data volume but lose physical meaning. Multi-fidelity
methods address measurement cost and label scarcity; they do not guarantee physical
consistency. If the dominant error comes from violating heat-transfer, melt-pool,
densification, or thermodynamic constraints, the next step is not only to combine
data sources but to constrain the learning problem itself. That point leads to physics-informed
and hybrid modeling.
3.4. Physics-Informed and Hybrid Learning for Physical Consistency
Transfer learning and multi-fidelity modeling improve how limited data are reused
and combined, but they do not guarantee physical consistency. A model may still predict
a plausible value while violating known process behavior, conservation laws, thermal
trends, or microstructure–property relationships. This risk is particularly relevant
in powder-based manufacturing because several controlling variables are only partially
observed. Powder-bed packing is heterogeneous, heat transfer is transient and geometry-dependent,
melt-pool behavior depends on the local powder state, and post-processing can modify
defects before final testing. A purely data-driven model may interpolate within a
narrow dataset but become unreliable outside the measured process window[13].
Physics-informed and hybrid learning address this problem by introducing prior physical
knowledge into the modeling process. In this review, physics-informed and hybrid learning
are used as broad umbrella terms to include PINNs and other physics-informed neural
architectures[25,
44,
54–
56], as well as broader hybrid models that use physical descriptors, simulation priors,
architecture constraints, or consistency checks[44,
45,
54–
56]. In practical terms, physical knowledge may enter before training through descriptors,
during training through constraints or architecture design, or after prediction through
consistency checks. In all cases, prior knowledge narrows, guides, or checks the learned
relationship instead of leaving the model to infer the full structure from data alone.
In powder-based manufacturing, the physical knowledge used in such models can take
several forms. LPBF thermal models often rely on heat-conduction or thermal-diffusion
equations, together with Gaussian, Goldak, or Rosenthal-type descriptions of a moving
heat source[24,
25,
45]. Melt-pool models may further involve conservation of mass, momentum, and energy
when fluid flow, recoil pressure, or free-surface effects are considered[25]. In powder metallurgy and post-processing, relevant priors may include Arrhenius-type
diffusion relations, sintering or densification kinetics, grain-growth models, phase-transformation
models, or precipitation-related models. For example, microstructure–property and
precipitation-related priors are particularly relevant in powder-metallurgy superalloy
design[50]. These models do not need to be solved fully inside every ML framework, but they
can guide feature design, constrain admissible predictions, or provide consistency
checks for learned trends[44,
54–
56]. Physics-informed Bayesian optimization can also incorporate simplified physical
priors into the search strategy when experiments or high-fidelity evaluations are
costly[57]. This distinction separates physics-informed and hybrid learning from the multi-fidelity
strategy discussed in Section 3.3. Multi-fidelity learning organizes evidence according
to cost, resolution, and measurement fidelity. Physics-informed learning organizes
the model according to physical constraints and mechanistic structure. The same simulation
can support either strategy, but its role is different. In a multi-fidelity workflow,
a simulation may act as a lower-cost data source. In a hybrid workflow, it may define
descriptors, constrain admissible outputs, or supply a mechanistic prior that shapes
the model response.
In powder-based additive manufacturing, many physics-informed studies focus on thermal
history and melt-pool evolution because these quantities link process parameters to
porosity, residual stress, solidification structure, and mechanical performance. Zhu
et al. developed a PINN framework for metal additive manufacturing by embedding governing
equations into models for temperature and melt-pool fluid dynamics[25]. Tiwari et al. later used a PINN for LPBF Inconel 718 by incorporating heat-transfer
constraints and a Goldak heat-source description, which represents the spatial distribution
of laser-induced heat input, into the learning framework[24]. Using a 347-sample multi-source dataset, their model reported mean absolute percentage
errors of 3.8%, 4.7%, 3.1%, and 1.9% for melt-pool width, melt-pool depth, peak temperature,
and relative density, respectively[24]. Ghungrad et al. followed a different route: an architecture-driven physics-informed
deep-learning model in which the network design was inspired by iterative transient-thermal
calculations for LPBF temperature prediction[45]. With 1000 points and an 80:20 split, the model reported a testing mean absolute
percentage error of about 2.8% and an R2 value of 0.936[45]. These examples show that “physics-informed” does not refer to a single algorithmic
form. Physical knowledge may enter through the loss function, the input representation,
the model architecture, the simulation prior, or checks applied to predicted outputs.
The relevant question is not whether a model is labeled physics-informed, but which
physical assumption is imposed and whether that assumption remains valid for the target
process regime.
Hybrid learning is broader than equation-constrained neural networks. In many powder-processing
studies, physical knowledge enters through descriptors rather than differential equations.
Examples include volumetric energy density, melt-pool geometry, cooling-rate estimates,
packing density, powder flowability, oxygen content, reuse count, precipitate volume
fraction, and diffusion-related features[6,
11,
50]. Descriptor-driven ML studies in alloy design also show that thermodynamic and electronic
descriptors can provide physically meaningful inputs for phase prediction, although
such descriptors still require experimental validation before being used for process
qualification[58]. Such descriptors can improve interpretability because they connect input variables
to powder, process, or microstructural mechanisms. A physically named descriptor,
however, is not physical validation. Volumetric energy density combines laser power,
scan speed, hatch spacing, and layer thickness into one scalar, but it can hide differences
in beam profile, scan strategy, powder absorption, shielding-gas flow, and local geometry.
The expected advantage of physics-informed learning is better behavior under sparse
data, but only along directions where the embedded physics remains valid. A heat-conduction
constraint can improve temperature prediction in a conduction-dominated LPBF regime,
where heat transport is mainly governed by thermal diffusion. A heat-conduction constraint
may be insufficient when non-conduction regimes such as keyhole-mode melting, vaporization-driven
spatter, denudation, or powder-bed disruption control the response. A sintering model
may transfer across related powder compacts when densification mechanisms are similar.
It may fail when liquid-phase sintering, swelling, abnormal grain growth, or phase
transformation dominates. Physics narrows the admissible solution space; it does not
identify the active mechanism by itself. Validation needs to test both numerical accuracy
and physical plausibility.
Predicted temperature fields should follow plausible thermal gradients and cooling
trends. Melt-pool dimensions should remain consistent with energy input and known
instability regimes. Predicted density or porosity should be checked against metallography,
CT, or other defect-sensitive measurements. Property predictions should be evaluated
against the microstructure and post-processing route that produced them. If a model
is claimed to be physically interpretable, its intermediate variables or learned trends
should be compared with known process–structure–property behavior rather than treated
as self-evident explanations. Interpretability is related to physical consistency,
but it is not the same thing. Hybrid descriptors, Shapley additive explanations (SHAP)
analysis, sensitivity analysis, and physically structured models can help identify
which variables influence a prediction, but they do not prove causality. Variable-importance
rankings may reflect oxygen content, reuse count, scan speed, or precipitate size
as important because those variables are correlated with hidden process conditions
in the dataset. Interpretation becomes more credible when powder lot, reuse history,
machine condition, geometry, post-processing route, and measurement method are recorded
well enough to reduce obvious confounding effects[6,
10,
11,
22]. Physics-informed and hybrid learning are most defensible when the relevant mechanism
is known well enough to constrain the model without excluding real behavior. The data
must also preserve links among inputs, intermediate process states, and final labels,
and validation must test both prediction accuracy and physical consistency. Under
these conditions, physical priors can improve sample efficiency and reduce nonphysical
predictions. Without them, adding physics terms or mechanistic descriptors may create
an appearance of rigor without improving reliability.
The three strategies discussed in Section 3 address different parts of the sparse-data
problem. Transfer learning relies on testable source–target similarity. Multi-fidelity
modeling relies on a meaningful hierarchy between lower-cost and higher-fidelity evidence.
Physics-informed learning relies on physical priors that remain valid for the target
regime. For powder-based manufacturing, none of these strategies removes the need
for lifecycle metadata. Physical constraints can guide learning, but they cannot compensate
for missing provenance, weak labels, unmeasured process shifts, or an incorrect assumption
about the governing mechanism.
4. CONCLUSIONS AND OUTLOOK
This review positioned ML in powder-based manufacturing as part of a powder-to-part
data chain, not as a collection of isolated algorithms. Its central distinction from
algorithm- or process-step-centered reviews is the emphasis on continuity among powder
history, process signatures, post-processing records, and validation labels across
the powder-to-part workflow. Current applications span feedstock optimization, powder
characterization, layer-quality monitoring, process-state detection, property prediction,
post-processing analysis, powder reuse assessment, and process optimization. Across
these areas, ML is most useful when it converts heterogeneous powder and process data
into testable predictions or decision support, and when those predictions remain traceable
to the powder state, processing route, sensor record, post-processing condition, and
final validation label.
The studies summarized in Table 1 show that reported model performance cannot be interpreted without the associated
data scale, validation design, target label, and deployment boundary. The limiting
factor is still the data structure behind the model. Many studies remain limited by
small datasets, uneven labels, narrow material or machine coverage, and incomplete
records of powder history and post-processing. High accuracy within a curated dataset
does not by itself show that a model will transfer to another powder lot, machine,
geometry, sensor setup, or heat-treatment route. For powder-based manufacturing, this
issue directly affects model interpretation and transferability. Powder reuse, oxidation,
particle-size evolution, recoating behavior, local thermal history, and post-processing
can all change the link between nominal input parameters and final properties. Without
those records, the model can learn a batch-specific or machine-specific correlation
while appearing to learn a general process–structure–property relationship.
Future work should move in two linked directions. The first is better data provenance.
Powder lot, reuse cycle, particle-size distribution, morphology, chemistry, oxygen
and moisture content, machine state, build location, sensor configuration, post-processing
route, and measurement method should be recorded in forms that can be linked across
the workflow. This type of digital thread is not just a data-management exercise;
it defines the boundary within which a model can be interpreted, reproduced, or transferred.
The second direction is more constrained learning. Transfer learning, multi-fidelity
modeling, and physics-informed or hybrid approaches are useful only when their assumptions
are explicit: source–target similarity for transfer learning, a meaningful fidelity
hierarchy for multi-fidelity modeling, and valid physical priors for physics-informed
learning.
ML is not a replacement for experimental judgment or qualification testing. Its predictions
are better viewed as structured evidence for choosing the next experiment, narrowing
a process window, flagging a possible defect, or prioritizing a candidate material
or post-processing route. A model that recommends a process condition still needs
validation against density, defect morphology, microstructure, mechanical performance,
and the relevant standard or application requirement. This is especially important
near the edge of the training domain, where changes in powder condition, defect mechanism,
or thermal regime can make a numerically plausible prediction physically unreliable.
Future ML systems for powder-based manufacturing will need to combine traceable lifecycle
data with models that have explicit applicability limits. Support for closed-loop
control, adaptive optimization, or digital-twin-oriented workflows requires more than
accurate regression or classification within a single dataset. They will need uncertainty
estimates, applicability-domain checks, physically meaningful features, and validation
protocols that account for shifts in powder lot, machine, geometry, sensor configuration,
and post-processing route. Progress in this field will therefore depend less on adding
more complex algorithms than on building datasets and models that preserve the links
among powder state, process evolution, intermediate signatures, and final performance.