
Chemometric-based Machine Learning for the Forensic Classification of Fire Debris
Key Takeaways
- Automating/semi-automating fatty-acid classification after ASTM E2881 can complement analyst review by improving throughput while preserving interpretability and traceability of the evidentiary basis for calls.
- Total ion chromatogram feature sets are susceptible to retention-time shifts and discard mass-spectral richness, allowing systematic artifacts to be learned when fixed scan segments encode misalignment.
Justin Miller-Schulze and Smith Purdum explain how machine learning and chemometrics can improve fire debris classification in forensic science applications.
LCGC International spoke to Justin Miller-Schulze and Smith Purdum about the development of a chemometric-based machine learning approach for the forensic classification of fire debris containing self-heating fatty acids following American Society for Testing and Materials (ASTM)E2881 analysis. The researchers aimed to create an automated or semi-automated workflow that could support human analysts by improving efficiency while maintaining transparency and interpretability. They initially investigated total ion chromatograms (TICs), but found that retention-time shifts and the loss of detailed mass-spectral information could negatively affect model performance. To address these limitations, three preprocessing workflows were developed and compared, ranging from relatively simple untargeted ratio processing to more detailed targeted approaches. Quantifier-to-qualifier ion ratios were selected as important features because they are reproducible and can be efficiently incorporated into automated workflows, according to the researchers.
The researchers also highlighted the importance of standardization for routine forensic implementation. Common retention-index (RI) standards and consistent GC–MS acquisition methods could improve robustness and allow laboratories to generate compatible datasets for machine learning. The researchers also see potential for extending the approach to more complex ignitable liquid residue analysis and suggested advanced techniques such as comprehensive two-dimensional gas chromatography (GC×GC) could help address co-elution, although they introduce additional data-processing challenges.
What was the rationale behind your paper Chemometric-based Machine Learning for the Forensic Classification of Fire Debris for the Presence of Self-heating Fatty Acids following Analysis by ASTM E2881?1
Classification of fire debris for fatty acids using ASTM E28812 traditionally relies heavily on the manual review and interpretation of the chromatographic and mass spectral data by a human analyst. We sought to investigate if an automated/semi-automated preprocessing and classification workflow could be developed that would complement a human analyst in a transparent and efficient manner.
What was the reasoning behind developing three different chemometric preprocessing workflows?
Our initial efforts attempted to use total ion chromatograms (TICs) for the input data to the machine learning models. However, TIC data obfuscates the richness of gas chromatograph mass spectrometry (GC–MS) data, in addition, slight to not-so-slight shifts in retention time (e.g., from column trimming) could influence the judgement of the model in a systematic manner. Using fixed scan segments as model features can cause relatively small retention time shifts to be encoded as artificial differences between otherwise similar samples, which inhibits training.
So, we knew we wanted a data pre-processing approach that classified molecular features in such a way that allowed for slight to not-so-slight retention time shifts and preserved the mass spectral richness of the data set to allow for better interpretation and classification by the model while at the same time being efficient. We used three different preprocessing workflows to see what the cost of more preprocessing would be - from relatively little (Untargeted Ratio) processing to relatively more preprocessing (Targeted) in terms of the classification performance.
Why did you choose quantifier-to-qualifier ion ratios as the main features for model training instead of raw chromatographic data?
The quantifier-to-qualifier ratios were easy to pull from standard preprocessing approaches and more reproducible for training of the machine learning algorithms. A library search method could have worked just as well for the Targeted approach, but ion ratios allowed for more automation in our Untargeted Ratio approach.
What chromatographic challenges had the greatest impact on feature extraction?
Relatively speaking, the chromatograms generated by ASTM E2881 from fire debris samples for fatty acid analysis are simple (as compared to something like ignitable liquid residue samples). That, in fact, was one of the goals of the work, to develop a workflow for a more straightforward data set that could potentially be transferred to a more complicated and demanding set of samples. So, co-elution and matrix effects were relatively minor issues in this data set.
Were there any chromatographic factors, such as retention time shifts or peak variability, that limited model performance?
Retention time variability was an issue when we initially used the TIC-preprocessing approach but the more refined, chemometric-based preprocessing methods applied in the paper addressed this for the most part. The chosen preprocessing methods were chosen precisely to eliminate the impact of retention time shifts or peak variability.
What additional chromatographic standardization would be needed before this workflow could be adopted routinely in forensic laboratories?
This was another goal of our work-to develop and compare a preprocessing workflow that could then be implemented in other laboratories to generate related datasets. These datasets could then be used to train the model and make it more and more proficient in its classifications.
Simple standardization approaches, like common retention index standards, would add robustness to the analytical and preprocessing workflows. With a common set of preprocessing methods, the generated data would be ready for input into a machine learning interface for classification. For routine adoption, however, the chemometric preprocessing would likely need to be incorporated into purpose-built software or processed off-site rather than requiring individual laboratories to reproduce the preprocessing manually. What is in the scope of capacity for different laboratories is adopting a singular GC–MS acquisition method, geared toward minimizing co-elution, which would ideally allow for an industry-wide pooling of training and reference data for use by machine learning models described in our paper.
Do you see this approach being extended to more complex GC–MS applications, such as ignitable liquid residue analysis, and what chromatographic challenges remain?
Absolutely. As described previously, this was always the intended next step of this work, and something we are actively working on. For the various ignitable liquid classes, an untargeted approach is theorized to very useful considering the dozens, if not hundreds of relevant molecules associated with these products. For more complex samples such as ILR (Ignitable Liquid Residue), more advanced chromatographic approaches may be needed to allow for maximum deconvolution of the sample data, such as comprehensive two-dimensional gas chromatography (GCxGC) due to co-elution. GCxGC systems produce highly complex data sets and the data interpretation is not a trivial endeavor by manual review. However, these two-dimensional data sets are likely to be ideal for machine learning pattern recognition approaches. While promising, GCxGC introduces challenges with time alignment which could produce artifacts for machine learning applications, similar to TIC preprocessing.
The majority of ignitable liquid residue analysis is essentially a visual pattern matching technique to a reference standard. The data presentations that human analysts generate are, by nature of human capability, ion profiles of only a few mass-to-charge ratios (m/z) ratios at a time. The limitation for humans is the amount of multi-dimensional information a person can practically interpret simultaneously. However, machines can process data much more quickly, process more data, and can theoretically discern more subtle patterns to achieve a goal. The recent advances of artificial intelligence (AI) in pattern matching indicates that the remaining challenge is in chromatographically preprocessing the data in a way that allows for classification by machine learning algorithms while maintaining a level of interpretability to allow for responsible human review.
Rather than asking a fire debris analyst to reduce a chemically complex sample to a handful of ion profiles, machine learning may ultimately allow us to evaluate a much larger portion of the chemical fingerprint while still presenting the evidence underlying that assessment back to the analyst. On the interface between chemistry and computer science, our goal is to standardize the chromatograms as much as possible to someday make the chemical fingerprints of ignitable liquids as amenable to automated database searching as human fingerprints. How much source-specific information may ultimately be recoverable from those chemical fingerprints remains an open question, and one that increasingly sophisticated chromatographic and machine learning approaches may allow the field to explore.
References
- Purdum, S.; Miller-Schulze, J. P. Chemometric-based Machine Learning for the Forensic Classification of Fire Debris for the Presence of Self-heating Fatty Acids Following Analysis by ASTM E2881. Forensic Science International: Reports 2026, 13, 100447. DOI:
https://doi.org/10.1016/j.fsir.2025.100447 . - Standard Test Method for Extraction and Derivatization of Vegetable Oils and Fats from Fire Debris and Liquid Samples with Analysis by Gas Chromatography-Mass Spectrometry.ASTM E-2881. ASTM International, 2022
https://compass.astm.org/content-access?contentCode=ASTM%7CE2881-18%7Cen-US
Related to this article








