6 Modeling Approaches That Make Complex Bioprocess Data Useful
By William Whitford, Oamaru BioSystems

A model is any representation of a real system that preserves those characteristics relevant to a particular purpose. Models have assumed many forms, from physical replicas and engineering drawings to computational simulations. Although they differ in construction and complexity, all models serve similar fundamental objectives: they simplify reality while maintaining the information necessary to describe, predict, or control the behavior of the system they represent.1
Bioprocess engineering employs many tools that are either types of models themselves or are model-enabled systems. This includes mechanistic, statistical, and AI models; design spaces, soft sensors, and controllers; as well as digital twins integrating several of these elements. Although they differ substantially in construction and purpose, all represent selected aspects of the manufacturing process for such objectives as process understanding, prediction, optimization, or control. These models increasingly draw upon real-time process data generated through process analytical technologies (PAT), advanced inline and online sensors, spectroscopy, imaging, and other analytical methods. Significantly, the variable, high-dimensional, nonlinear, and higher-order relationships governing biological systems are very difficult to specify accurately in advance.2
Artificial intelligence (AI) has recently contributed powerful, objective-driven dimensionality reduction that compresses diverse and complex process measurements while preserving information relevant to prediction and decision-making. When integrated with quality by design-defined spaces, such activity can powerfully support process evaluation, trajectory prediction, and model-predictive control — enabling more adaptive operation of biomanufacturing systems.3 Rather than replacing the traditional concept of a model, AI extends it by enabling aspects of the model to be learned directly from increasingly complex and multimodal data.
The Bioprocessing Data Challenge
Bioprocess data have historically been analyzed using tools such as Nelson rules, partial least-squares models, and response-surface models. These approaches have served the industry well, particularly when process data were limited and of relatively low dimensionality. However, the volume and complexity of available data have grown with the introduction of new analytical modalities and more frequent process monitoring. Modern bioprocess data now span several overlapping categories.
Comprehensive historical data from completed batches can be used to, for example, identify recurring patterns, whereas such real-time data as continuously recorded pH and dissolved-oxygen measurements can support monitoring and control during operation. These real-time data are available during the process through sensors or automated analyzers, while offline data include such measurements as metabolite concentrations or product quality attributes obtained from samples analyzed in the laboratory.
Bioprocess data may also be classified or characterized along other dimensions. For example, labeled data supports the association of process observations with a known outcome, such as product titer or glycosylation. Whereas unlabeled data, including large collections of sensor measurements without corresponding outcome assignments, may still contain patterns that robustly characterize process behavior. Structured data are organized into predefined fields, as in historian tables containing timestamps, process parameters, and batch identifiers. Unstructured data include microscopy images, operator observations, and free-text batch record entries. Experimental data are generated under deliberately varied conditions, such as a design of experiments study of temperature and pH, whereas operational data arise during prescribed routine manufacturing. Data may also be measured directly, such as temperature from a glass combination electrode, or inferred from other measurements, such as viable cell density estimated from a calibrated capacitance sensor. A single data source may belong to several of these categories.
The relationships within these data present additional challenges. Examples of nonlinearity include substrate inhibition and the relationships between environmental conditions and cell growth. Higher-order interactions arise when the effect of one variable depends on combinations of others, such as the combined effects of temperature, pH, and nutrient availability on product quality. High dimensionality occurs in data generated by, e.g., spectroscopy, cell imaging, and multi-omics analyses. Discontinuities may result from discrete process events, including inoculation, feed initiation, and media exchange. Together, these data characteristics can make bioprocess relationships difficult to represent using simple linear or low-dimensional models 4,5
AI algorithms can reduce the dimensionality of such data more faithfully by learning compact representations that preserve nonlinear relationships, higher-order interactions, and other multivariate features relevant to prediction or control.6,7 Each of the “models” below offers a different approach for compressing complex data into the kind of representation that makes it useful to a particular objective.
Classical Process Models
Classical bioprocess engineering represents manufacturing operations using engineer-defined variables whose physical or process meaning is specified in advance.5 This representation may be implemented through mechanistic, statistical, empirical, or hybrid models. Depending on the structure of each, these models variously struggle to accommodate high-dimensional, heterogeneous, discontinuous, or strongly nonlinear data.
Mechanistic models, such as Monod- or Michaelis–Menten-type models, rely on rather simplified assumptions about underlying physical and biological mechanisms. Their accuracy may be quite limited when important process mechanisms are unknown or change during processing. For example, a Monod-based cell-growth model may accurately predict a CHO culture during exponential growth but fail after a change in cell physiology.
Statistical and empirical models describe relationships found in their training data but may not functionally reflect the underlying biological mechanisms. Consequently, they may perform poorly when extrapolated beyond, for example, the materials or operating ranges that were used in their development. For example, a PLS model trained to predict viable cell density from Raman spectra may lose accuracy after a change in media formulation. An important strength of representation-learning AI models is their ability to learn, not only relationships, but also a dynamic internal representation, or latent space, using mechanistic knowledge together with historical and newly generated empirical data.8
Design Space Models
A design space is a multidimensional representation of values defining acceptable operation. This region of acceptable operation is defined by material attributes and process parameters demonstrated to assure product quality. Its value is not only to distinguish acceptable operating conditions, but to provide a scientifically justified basis for process adjustment. A bioreactor design space commonly includes such parameters as temperature, pH, and metabolite concentrations, while the actual number of values considered varies. Acceptable performance includes such process outcomes as product titer, aggregation, and glycosylation. A design space can be viewed as a representational model because it maps selected process inputs to regions associated with experimentally established acceptable production performance.9
Soft Sensing Models
Classical soft sensor modeling tools include mass balance calculations, partial least-squares models, and Kalman filters. In biomanufacturing, they can be used to estimate such quantities as viable cell density and metabolite concentrations from measured process variables and online sensor signals. These estimates support process monitoring when direct measurements are unavailable or too costly for routine use. More recently, AI models using spectroscopic measurements or images have been examined as soft sensors. A soft sensor can be considered a model because it represents the relationship between observable process information and an unmeasured quantity of interest, enabling that quantity to be inferred rather than measured directly.10
Latent Space Models
AI is changing process modeling in several distinct ways. It can faithfully reduce the dimensionality of complex process data while preserving features relevant to prediction or control. Through representation learning, an AI algorithm can identify combinations of variables that summarize dominant patterns without requiring specification in advance. A latent space is a learned coordinate system whose dimensions represent underlying factors inferred from high-dimensional process observations. The resulting latent variables preserve the information most relevant to the prescribed objective. Each coordinate may capture a recurring pattern associated with, for example in biomanufacturing, cell behavior or product formation — although it may not directly correspond to a single measurable quantity. By mapping multimodal data into this reduced space, latent space models can facilitate batch comparison, anomaly detection, and process prediction. They can be considered a “model” because the learned representation preserves selected characteristics of the process for defined analytical or operational purposes.11
Latent State Models
Using a learned representation of process data, AI algorithms can estimate the current process state as a “latent state” — an underlying condition that cannot be observed directly. Unlike a latent space, a latent state model incorporates the sequence and history of process measurements to include how the system has evolved over time. This state estimation can help identify, for example, transitions between growth, production, and decline. It is also valuable for detecting departures from expected trajectories, predicting future behavior, and supporting feedback control. Measured process variables provide the observations from which the latent state is inferred, while manipulated variables will influence how that state changes. One might consider a latent state estimation as related to soft sensing, but the concepts differ in emphasis. A soft sensor generally estimates a specific unmeasured quantity, whereas a latent state provides a comprehensive representation of the overall biological condition of the process.12
Digital Twin Models
A digital twin or digital shadow is a comprehensive virtual representation of a physical system or device that is continually updated using data from its physical counterpart. In biomanufacturing, it may integrate mechanistic and statistical models, soft sensors, latent state estimators, equipment models, and process data within a common architecture. Digital twins and digital shadows are proving valuable for process prediction and optimization because they can simulate how a process may respond to changing conditions before an action is implemented. They support process controllers by estimating the current state, forecasting future trajectories, and identifying actions that satisfy established operating constraints (operating space). A digital shadow differs from a twin as it contains a more limited architecture in which information flows only from the physical process to its virtual representation.13 Digital twins are particularly well suited to managing and controlling complex systems such as upstream and downstream bioprocesses, where an added layer of regulatory requirements must also be respected. In this sense, a digital twin can be understood as a cyber-system built from a coordinated set of AI models able to interpret and interact with the physical process in real time, adjusting operation according to the continuous feedback provided by biological and process environment conditions.6,13

Conclusion
Bioprocess modeling is evolving in response to the increasing volume, diversity, and complexity of biomanufacturing data. Classical process models remain valuable because, as they encode engineering knowledge and provide reliable predictions with relatively limited data, they can be sufficient for particular applications. Some of the approaches defined above extend this foundation by modeling acceptable operation and estimating quantities that cannot be measured directly. They simplify complex process data while preserving information needed to predict or control its behavior.14
AI expands this modeling continuum by allowing both process relationships and the consequences of their underlying mechanisms to be learned from data. Latent space models condense high-dimensional, multimodal observations into more informative representations, while latent state models use the history of those observations to estimate the current and future biological consequences. These capabilities are particularly relevant to bioprocesses, where nonlinear relationships, interacting variables, and changing cellular states are impossible to specify accurately in advance.
Moving these individual AI models into routine biomanufacturing operation raises a further consideration: scale-up. A biomanufacturing process typically requires multiple AI models operating in parallel, often one model per critical quality attribute, together with additional models that continuously monitor the primary models for drift over time. Deploying and maintaining this larger population of models in production, while sustaining the compliance evidence expected by regulators, calls for dedicated platforms. The demand for this transfer of technology has become so ubiquitous that a whole new industry has risen up to support it.
Personally, I've worked with the founders at Aizon and seen its platform for scaling single validated models into coordinated production models up close. Other firms offer similar products with subtle to substantial differences depending on capability needs. At the orchestration level, Modersys, formerly SmartFactory Rx, emphasizes alarm management and maintaining process parameters within the design space. Quartic's systems synthesize multiple legacy PAT signals focusing on optimization and predictive maintenance. Some solutions are more specialized, for example, Ansys' systems for bioprocessing are specific to fluid dynamics, combining modeling and artificial intelligence to predict flow patterns and gas distribution in bioreactors.15
The coordinating infrastructures described here integrate mechanistic knowledge, empirical relationships, learned representations, state estimates, and real-time data according to the requirements of the application. The future of bioprocess development and manufacturing will therefore depend not on choosing between classical and AI-based models, but on integrating their complementary strengths to build sophisticated in silico systems that are accurate, maintainable, explainable, and fit for purpose.
References:
- Mario Stassen, Matt Schmucki, Francisco Valero, Toni Manzano, From Traditional Statistics to Adaptive Multivariate Models: Exploring the Role of AI in Modern Pharmaceutical Manufacturing, PDA Journal of Pharmaceutical Science and Technology Jul 2026, pdajpst.2026-000011.1; DOI: 10.5731/pdajpst.2026-000011.1
- Santo Motta and Francesco Pappalardo, “Mathematical Modeling of Biological Systems,” Briefings in Bioinformatics 14, 411–422 (2013).
https://doi.org/10.1093/bib/bbs061 - Toni Manzano, William Whitford, Chapter 4 - AI applications for multivariate control in drug manufacturing, Editor(s): Anil Philip, Aliasgar Shahiwala, Mamoon Rashid, Md. Faiyazuddin, A Handbook of Artificial Intelligence in Drug Delivery, Academic Press, 2023, Pages 55-82, ISBN 9780323899253, https://doi.org/10.1016/B978-0-323-89925-3.00023-X.
- Matthew Banner, Haneen Alosert, Christopher Spencer, Matthew Cheeks, Suzanne S. Farid, Michael Thomas, and Stephen Goldrick, “A Decade in Review: Use of Data Analytics within the Biopharmaceutical Sector,” Current Opinion in Chemical Engineering 34, 100758 (2021).
https://doi.org/10.1016/j.coche.2021.100758 - Bailey, J.E., “Mathematical Modeling and Analysis in Biochemical Engineering: Past Accomplishments and Future Opportunities.”
https://doi.org/10.1021/bp9701269 - Toni Manzano and William Whitford, “Artificial Intelligence Empowering Process Analytical Technology and Continued Process Verification in Biotechnology.”
https://doi.org/10.1089/genbio.2024.0041 - Geoffrey E. Hinton and Ruslan R. Salakhutdinov, “Reducing the Dimensionality of Data with Neural Networks,” Science 313, 504–507 (2006).
https://doi.org/10.1126/science.1127647\ - Hinton and Salakhutdinov, “Reducing the Dimensionality of Data with Neural Networks.”
https://doi.org/10.1126/science.1127647 - Anurag S. Rathore and Helen Winkle, “Quality by Design for Biopharmaceuticals,” Nature Biotechnology 27, 26–34 (2009).
https://doi.org/10.1038/nbt0109-26 - Reiner Luttmann, Daniel G. Bracewell, Gesine Cornelissen, Krist V. Gernaey, Jarka Glassey, Volker C. Hass, Christian Kaiser, Christian Preusse, Gerald Striedner, and Carl-Fredrik Mandenius, “Soft Sensors in Bioprocessing: A Status Report and Recommendations”
https://doi.org/10.1002/biot.201100506 - John F. MacGregor, Honglu Yu, Salvador García-Muñoz, and Jesus Flores-Cerrillo, “Data-Based Latent Variable Methods for Process Analysis, Monitoring and Control,” Computers & Chemical Engineering 29, 1217–1223 (2005).
https://doi.org/10.1016/j.compchemeng.2005.02.0071 - Harini Narayanan, Lars Behle, Martin F. Luna, Michael Sokolov, Gonzalo Guillén-Gosálbez, Massimo Morbidelli, and Alessandro Butté, “Hybrid-EKF: Hybrid Model Coupled with Extended Kalman Filter for Real-Time Monitoring and Control of Mammalian Cell Culture,” Biotechnology and Bioengineering 117, 2703–2714 (2020).
https://doi.org/10.1002/bit.27437 - William G. Whitford and Toni Manzano, “AI-Enabled Digital Twins in Biopharmaceutical Manufacturing,” BioProcess International 21(7–8), 22–27 (2023).
https://www.bioprocessintl.com/information-technology/ai-enabled-digital-twins-in-biopharmaceutical-manufacturing - Smiatek, Jung, and Bluhmki, “Towards a Digital Bioprocess Replica.”
https://doi.org/10.1016/j.tibtech.2020.05.008 - Singh, Jiménez del Val, Glassey, Kavousi, "Integration Approaches to Model Bioreactor Hydrodynamics and Cellular Kinetics for Advancing Bioprocess Optimisation," Bioengineering (Basel) 2024 May 27;11(6):546.
https://doi.org/10.3390/bioengineering11060546
About The Author:
William Whitford is founder and principal at Oamaru BioSystems, a consulting company. He has over 20 years of experience leading biotechnology product and process development teams. Past roles include leadership positions at Arcadis, Cytiva, and Thermo Fisher Scientific.