An Analysis of Artificial Intelligence-Based Predictive Models for Improving Data Security and Decision-Making in Cloud Computing

Main Article Content

Dr. Rajesh Kumar Sharma

Abstract

Cloud services provide vast amounts of security telemetry, but the scale of alerts and the asymmetric impact of missed threats impede quick action. In this work we explore the ability of supervised prediction models to classify botnet flows from benign flows in a public AWS-hosted testbed and propose a human-in-the-loop alarm triage architecture. A derivative of CSE-CIC-IDS2018 with 48,000 records was created and split according to the provided training and testing blocks of 24,000 records each with 500 bot and 23,500 benign flows. We removed the label and the provenance index to be left with 78 numeric flow predictors. Median imputation and model specific scaling were fitted on training data We evaluated majority-class, class-weighted logistic-regression, random-forest, and histogram-gradient-boosting classifiers on the original test labels. With a fixed threshold of 0.50, histogram boosting yielded 497 true positives, 3 false negatives, 2 false positives and 23,498 true negatives. The accuracy was 0.9960, the recall was 0.9940, the F1 was 0.9950, and the average precision was 0.9978. This model achieved an F1 of 0.9927 and average accuracy of 0.9961 by removing 874 test records with an identical feature-and-label signature of training records. These findings indicate a high distinction for this limited bot vs. benign benchmark. The triage bands provided indicate how scores may be used to organise analyst evaluation, although neither decision quality nor protection of actual cloud assets was tested. Still need External validation, calibration and future incident-response assessment.

Article Details

How to Cite
Dr. Rajesh Kumar Sharma. (2026). An Analysis of Artificial Intelligence-Based Predictive Models for Improving Data Security and Decision-Making in Cloud Computing. International Journal of Advanced Research and Multidisciplinary Trends (IJARMT), 3(3), 1155–1170. https://doi.org/10.65578/ijarmt.v3.i3.1314
Section
Articles

References

Breiman, L. (2001). Random forests. Machine Learning, 45, 5–32. https://doi.org/10.1023/A:1010933404324

Buczak, A. L., & Guven, E. (2016). A survey of data mining and machine learning methods for cyber security intrusion detection. IEEE Communications Surveys & Tutorials, 18(2), 1153–1176. https://doi.org/10.1109/COMST.2015.2494502

Canadian Institute for Cybersecurity. (2018). CSE-CIC-IDS2018 on AWS [Data set]. University of New Brunswick. https://www.unb.ca/cic/datasets/ids-2018.html

Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953

Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10.1145/2939672.2939785

Cloud Security Alliance. (2024). Top threats to cloud computing 2024. https://cloudsecurityalliance.org/artifacts/top-threats-to-cloud-computing-2024

Davis, J., & Goadrich, M. (2006). The relationship between precision-recall and ROC curves. In Proceedings of the 23rd International Conference on Machine Learning (pp. 233–240). ACM. https://doi.org/10.1145/1143844.1143874

Fawcett, T. (2006). An introduction to ROC analysis. Pattern Recognition Letters, 27(8), 861–874. https://doi.org/10.1016/j.patrec.2005.10.010

Friedman, J. H. (2001). Greedy function approximation: A gradient boosting machine. The Annals of Statistics, 29(5), 1189–1232. https://doi.org/10.1214/aos/1013203451

Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (Vol. 70, pp. 1321–1330). PMLR. https://proceedings.mlr.press/v70/guo17a.html

He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering, 21(9), 1263–1284. https://doi.org/10.1109/TKDE.2008.239

Jansen, W., & Grance, T. (2011). Guidelines on security and privacy in public cloud computing (NIST SP 800-144). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.800-144

Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems (Vol. 30, pp. 3146–3154). https://papers.nips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree

Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30). https://proceedings.neurips.cc/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html

Mell, P., & Grance, T. (2011). The NIST definition of cloud computing (NIST SP 800-145). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.800-145

Moustafa, N., & Slay, J. (2015). UNSW-NB15: A comprehensive data set for network intrusion detection systems. In 2015 Military Communications and Information Systems Conference (MilCIS) (pp. 1–6). IEEE. https://doi.org/10.1109/MilCIS.2015.7348942

National Institute of Standards and Technology. (2011). Information security continuous monitoring (ISCM) for federal information systems and organizations (NIST SP 800-137). https://doi.org/10.6028/NIST.SP.800-137

National Institute of Standards and Technology. (2020). Security and privacy controls for information systems and organizations (NIST SP 800-53 Rev. 5). https://doi.org/10.6028/NIST.SP.800-53r5

National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). https://doi.org/10.6028/NIST.AI.100-1

National Institute of Standards and Technology. (2025). Incident response recommendations and considerations for cybersecurity risk management: A CSF 2.0 community profile (NIST SP 800-61 Rev. 3). https://doi.org/10.6028/NIST.SP.800-61r3

Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12, 2825–2830. https://www.jmlr.org/papers/v12/pedregosa11a.html

rayborg. (n.d.). Download and use the public bundle: CSE-CIC-IDS2018 bot-versus-benign preserved-ratio CSV [Data set]. DCTABGAN IDS benchmark datasets. https://rayborg.github.io/dctabgan-ids-benchmark-datasets/download.html

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should I trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1135–1144). ACM. https://doi.org/10.1145/2939672.2939778

Saito, T., & Rehmsmeier, M. (2015). The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLOS ONE, 10(3), e0118432. https://doi.org/10.1371/journal.pone.0118432

Sharafaldin, I., Habibi Lashkari, A., & Ghorbani, A. A. (2018). Toward generating a new intrusion detection dataset and intrusion traffic characterization. In Proceedings of the 4th International Conference on Information Systems Security and Privacy (pp. 108–116). SCITEPRESS. https://doi.org/10.5220/0006639801080116

Sommer, R., & Paxson, V. (2010). Outside the closed world: On using machine learning for network intrusion detection. In 2010 IEEE Symposium on Security and Privacy (pp. 305–316). IEEE. https://doi.org/10.1109/SP.2010.25

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.