Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach
Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach
基于决策树与 K-Means 的食用油拉曼光谱分析:一种物理信息驱动的 AI 方法
Abstract: Authentication of edible oils in processed foods is important for food quality, fraud prevention, and regulatory compliance. This study establishes an integrated Raman spectroscopy and machine-learning framework that links intrinsic spectral organization, interpretable classification, and Physics-Informed Artificial Intelligence (PI-AI).
摘要: 加工食品中食用油的真伪鉴别对于食品质量、防范欺诈及合规监管至关重要。本研究建立了一个集拉曼光谱与机器学习于一体的框架,将内在的光谱组织结构、可解释的分类方法以及物理信息驱动的人工智能(PI-AI)有机结合。
Five edible oils were investigated in pure form and within a fried-potato-chip matrix using t-SNE, K-means clustering, Decision Trees, and Non-Negative Least Squares (NNLS)-based spectral decomposition. Unsupervised analyses revealed substantially stronger class organization and separability in pure oils, whereas food-matrix effects introduced pronounced spectral overlap.
研究人员利用 t-SNE、K-means 聚类、决策树以及基于非负最小二乘法(NNLS)的光谱分解技术,对纯食用油及炸薯片基质中的食用油进行了分析。无监督分析显示,纯油样本具有更强的类别组织性和可分性,而食品基质效应则导致了显著的光谱重叠。
Decision Trees achieved 100% classification accuracy for pure oils using only four Raman variables from the original 1866-feature spectral space. These four variables, consistently identified by both pre-pruned and post-pruned models, represented only approximately 0.21% of the available spectral information while retaining perfect test-set performance.
决策树模型仅利用原始 1866 个光谱特征空间中的 4 个拉曼变量,就实现了纯油样本 100% 的分类准确率。这 4 个变量在预剪枝和后剪枝模型中均被一致识别,它们仅占可用光谱信息的约 0.21%,却依然保持了完美的测试集表现。
For matrix-containing samples, NNLS-based PI-AI spectral decomposition substantially improved classification by separating oil-related signatures from paper and potato contributions. Optimized post-pruned models achieved accuracies of 86.4% and 85.4% for paper-subtracted and paper-plus-potato-subtracted datasets, respectively, while reducing the number of important Raman variables to only five and four.
对于含有基质的样本,基于 NNLS 的 PI-AI 光谱分解技术通过将油类特征信号与纸张及马铃薯成分分离,显著提升了分类效果。优化后的后剪枝模型在去除纸张干扰和去除纸张及马铃薯干扰的数据集上,分别达到了 86.4% 和 85.4% 的准确率,同时将关键拉曼变量的数量分别缩减至 5 个和 4 个。
The compact four-feature representation further reduced the data footprint by 99.44% without loss of classification accuracy. Collectively, these findings demonstrate that accurate Raman-based oil identification can be achieved through physically meaningful, highly compact, and interpretable spectral representations, providing a promising foundation for Frugal AI, Edge AI, portable sensing, and embedded food-quality monitoring.
这种紧凑的四特征表示在不损失分类准确率的前提下,将数据占用空间进一步减少了 99.44%。总而言之,这些研究结果表明,通过具有物理意义、高度紧凑且可解释的光谱表示,可以实现精确的拉曼油品识别,这为节俭 AI(Frugal AI)、边缘 AI、便携式传感及嵌入式食品质量监测奠定了坚实基础。