About Me
Hello, I am Mohammad Saef Ullah Miah. I am an Associate Professor at the department of Computer Science, AIUB. I have obtained my PhD from Universiti Malaysia Pahang Al-Sultan Abdullah, and have completed my Master's and Bachelor's degree from AIUB. I have hands-on experience in Software Development, Machine Learning, Deep Learning, Natural Language Processing, Data Science, Blockchain and project management. Prior to that, I completed my Higher Secondary education from Rajuk Uttara Model College and Secondary schooling from Ideal School and College.
Apart from my routine work, I enjoy managing open source projects and exploring new technological advancements and gadgets. My research interests lie in various areas including Machine Learning, Natural Language Processing, Data and Text Mining, Knowledge Management, Material Informatics, and Blockchain based decentralized applications. I am passionate about applying these technologies to solve real-world problems and contribute to the development of innovative solutions.
Journal Publications
BRNLTK: A Python Toolkit for Processing Bangla Regional Dialects
Abstract
Bangla is a morphologically rich, low-resource language whose regional dialects suffer from an even greater scarcity of tools, corpora, and annotated datasets. This paper introduces brnltk, the first modular Python toolkit designed for processing Bangla regional dialects. It integrates rule-based, statistical, and neural components for dialect translation, POS tagging, stemming, tokenization, and semantic similarity computation across five regions: Mymensingh, Barishal, Chattogram, Noakhali, and Sylhet. POS tagging reaches up to 96% accuracy, exceeding a general-purpose Bangla baseline (50.8%); dialect translation attains chrF 25.0–54.1, BLEU 1.4–13.4, METEOR 11.2–37.5, WER up to 89.4%, and CER up to 64.3% across dialect pairs. Training-free for end users and built on human-verified dialectal mappings, brnltk requires no annotated corpora or fine-tuning.
A Comprehensive Dataset for Postpartum Depression Prediction Among Bangladeshi Mothers: Sociodemographic, Obstetric, and Psychosocial Factors with Class Imbalance Insights
Abstract
Postpartum depression (PPD) is currently a serious mental disorder that occurs after childbirth. It affects the well-being of both the mother and the child. However, early screening is still limited in low- and middle-income countries like Bangladesh due to socio-economic challenges and social stigma. This paper presents a comprehensive and multidimensional dataset consisting of 766 participants to facilitate data-driven PPD research. The participating women were between the ages of 18-41 years. The data were collected through a cross-sectional study using face-to-face interviews and hospital records at LABAID Diagnostics, Pabna, Bangladesh, between March and June 2025. The dataset integrates 51 variables, including relevant factors such as sociodemographic factors, obstetric history, neonatal details, and psychosocial indicators. Maternal health was assessed using standardized screening tools: PHQ-2 and PHQ-9 for symptoms before and during pregnancy, and Edinburgh Postpartum Depression Scale (EPDS) for postpartum symptoms. Results indicated that the prevalence of depression during pregnancy among the participants was 40.47%. A key feature of this dataset is the clear reporting of class imbalance, which can be reused for statistical analysis, epidemiological studies, and validation of depression screening tools, making it a valuable resource for building robust machine learning (ML) and interpretable AI (XAI) models for early risk prediction. This study fills an important gap in publicly available, high-quality maternal health information in the South Asian context. It also provides the basis for evidence-based interventions and targeted support for at-risk mothers.
Predicting Climate Change Impacts on Temperature and Rainfall in Bangladesh Using Ensemble Learning: A Statistical Yearbook Analysis
Abstract
Climate change in Bangladesh is causing unstable temperatures and rainfall, increasing the risk of floods, droughts, and agricultural disruption. Previous studies missed integrating temperature and rainfall data with feature selection and Explainable AI (XAI). This paper introduces a novel approach that integrates machine learning (ML) and deep learning (DL) methods with feature selection and XAI to improve climate impact predictions across the region. Using a comprehensive dataset from the Bangladesh Statistical Yearbook (1981–2023), our feature analysis highlights that key months such as April, February, and June strongly influence temperature and rainfall predictions, while location plays a greater role in shaping rainfall patterns than temperature. Our proposed ensemble model outperformed traditional and previous models across three climate datasets, achieving \({\textrm{R}}^{2}\) scores of 1.00, 0.79, and 0.99 for maximum temperature, minimum temperature, and rainfall, respectively. It also recorded low MAE values of 0.08, 0.39, and 3.46, and RMSE values of 0.10, 0.50, and 4.20 correspondingly. XAI techniques like LIME improved prediction interpretability by revealing key factors influencing the model’s decisions. Critical influences included lower March and July temperatures for maximum temperature predictions, while higher January and February temperatures positively affected the model. For rainfall, significant contributions from July and September rainfall were observed, while lower rainfall in April and December negatively impacted predictions. These insights support better agriculture and water management by identifying key climate drivers. Future work can enhance the model by adding more meteorological variables and considering large-scale climate patterns for deeper regional insights.
CompoundDenseNet: a novel approach for accurate recognition of Bangla handwritten compound characters
Abstract
Bangla, one of the most widely spoken languages in the world, presents major challenges in handwritten character recognition because of its complex compound characters with intricate shapes, diverse writing styles, and structural similarities. These features make Bangla a representative example of complex scripts that remain difficult for conventional Optical Character Recognition (OCR) systems. This study focuses on improving the recognition of Bangla handwritten compound characters using a modified DenseNet architecture named CompoundDenseNet. The architecture enhances feature extraction and reuse to better capture the visual variations and fine structural details that existing models often struggle to handle. Its performance was evaluated on three benchmark datasets, BanglaLekha Isolated, Ekush, and CMATERdb, achieving recognition accuracies of 98.5%, 98%, and 96.2% respectively, surpassing previously reported methods. Misclassification analysis using a confusion matrix revealed that the Adam optimizer produced the most stable and accurate results with faster convergence compared to other optimizers tested. While the results demonstrate significant progress, the study also highlights the need for larger and more diverse datasets. Overall, CompoundDenseNet contributes to advancing Bangla handwritten compound character recognition and has the potential to enhance real-world applications such as education, legal documentation, and digital accessibility in Bangla language technologies.
A review of fairness challenges in natural language processing
Abstract
Natural Language Processing (NLP) systems are increasingly deployed in high-stakes systems including healthcare, education, recruitment, and law enforcement, yet they have frequently coded and magnified biases that undercut their system’s fairness and trust. This review synthesizes and critically analyzes 121 studies that were published in the year 2014 and up to date that address bias in NLP. We present a novel taxonomy of 18 bias types, such as previously underexplored categories like geographic, disability, and annotation bias, and project them onto the NLP lifecycle, taking data as the starting point to deployment. Four key detection paradigms are examined (statistical, model-probing, benchmark-based, and human-centric), alongside mitigation strategies at the data, model, and post-processing levels. Unlike prior surveys, this study offers a lifecycle-aware framework that connects bias origins, detection methods, and mitigation practices, while focusing on persistent challenges such as intersectionality, generalization, and fairness–performance trade-offs in large language models (LLMs). We argue that achieving fairness in NLP requires not only technical interventions but also socio-technical approaches that integrate community participation, transparency, and governance. Offering a structured, critical, and forward-looking synthesis, this work contributes a roadmap for building transparent, equitable, and socially responsible NLP systems.
An efficient dual path deep learning framework for COVID-19 classification using lung CT scans with explainable AI
Abstract
While the global burden of COVID-19 has eased due to widespread vaccination and public health efforts, the virus has not been eradicated. New variants continue to emerge, and localized outbreaks remain a concern, particularly in regions with limited healthcare resources. This highlights the ongoing need for rapid, accurate, and scalable diagnostic tools. In this study, a comprehensive deep learning framework for detecting COVID-19 from lung CT scans is presented, aimed at improving diagnostic reliability and computational efficiency. An extensive and diverse CT dataset was curated by combining images from nine publicly available datasets, with a total of 25,408 samples in COVID-19 and normal classes. Multiple state-of-the-art convolutional neural networks (CNNs) and vision transformer models were fine-tuned and evaluated under consistent conditions to build a strong performance benchmark. Based on these findings, a new lightweight parallel model was developed, combining a custom CNN and a pretrained backbone. Both networks process the input image independently, and their extracted features are fused at the final stage for classification. The proposed model demonstrated higher accuracy (97.46%) compared to other models tested in this study, while maintaining low computational complexity. Additionally, explainable AI techniques, including Grad-CAM and LIME, were employed to provide visual interpretations of the model’s predictions.
HyOPTEnsemble: custom-weighted soft voting hyperparameter optimization ensemble model, explainable-AI for predicting mental state among university students
Abstract
The emergence of mental health issues among university students has a detrimental impact on their academic and job performance, leading to increased levels of anxiety, dropout mentality, and even suicide. Male students are at higher risk compared to their female counterparts. Early detection of mental health issues is crucial for effective treatment. While previous research has focused primarily on specific departments or universities, few studies have examined data from multiple institutions. This study proposes a novel hypertuned ensemble machine learning model that demonstrates its effectiveness in predicting the mental state of university students based on their past two weeks’ experiences. The dataset was collected through an online survey conducted over a period of one month, involving 50 universities, with a total of 400 responses from undergraduate and graduate students. Of these, 67.1% were male, 31.7% were female, and 1.2% declined to disclose their gender identities. To address the data imbalance issue, a hybrid synthetic minority oversampling method, SMOTE-ENN, is employed. Initially, single machine learning models were utilized, including Decision Tree (DT), Naive Bayes (NB), Ada-Boost, Bagging, Random Forest (RF), and Modified Random Forest. In addition, a grid search was performed for hyperparameter optimization. However, the ensemble model with 5-fold cross-validation achieved the highest accuracy of 0.977, along with outstanding performance metrics of precision, recall, and f1 score, at 0.983, 0.977, and 0.979, respectively. Furthermore, Cohen’s kappa coefficient and Matthews correlation coefficient were utilized to measure reliability. The proposed model demonstrates the ability to predict the mental state of university students.
Explainable AI framework for improved Thalassemia mental health classification and feature selection
Abstract
Mental health challenges in Thalassemia patients are often overlooked, despite their significant impact on quality of life. Traditional statistical and machine learning approaches often fail to capture the complex, nonlinear relationships between psychosocial and clinical variables, limiting both accuracy and interpretability. To overcome this, we propose the Adaptive Multi-Stage Ensemble with Dynamic Feature Interaction (AMSE-DFI)-a novel feature selection framework that dynamically integrates mutual information, ensemble learning, and graph attention mechanisms to capture intricate feature interdependencies often missed by traditional approaches. Using the SF-36 health survey data from 356 Bangladeshi patients, AMSE-DFI effectively identified key predictors such as total SF score, role emotional, and physical health summary, which collectively reflect both physical and psychological well-being. The model outperformed conventional approaches, showing strong predictive reliability and robust generalization, with SMOTE effectively addressing class imbalance in the clinical data. Importantly, Local Interpretable Model-Agnostic Explanations (LIME) based explainable AI offered clear, interpretable insights into how key features affect individual patient outcomes, making the model more understandable and actionable for clinicians. This framework provides a practical, transparent tool to support early detection and personalized management of mental health challenges in Thalassemia care.
A clinical dataset on type-2 diabetes including demographic, anthropometric, and biochemical parameters from Bangladesh
Abstract
Type-2 diabetes is a major public health concern in Bangladesh, and this dataset provides 1065 curated patient records with demographic, anthropometric, and clinical variables relevant to its assessment. The data were collected during routine clinical visits and recorded by trained staff, with checks to ensure accuracy and completeness. It includes basic details like age, pregnancy count, body mass index, and skin-fold thickness; vital signs such as blood pressure; lab results related to blood sugar (fasting glucose and insulin); the Diabetes Pedigree Function; and a simple yes/no label for Type-2 diabetes. A few values are missing for diastolic blood pressure and skin-fold thickness, so users should handle these carefully. Since the data are cross-sectional and come from patients seeking care, there are more diabetic cases (840) than non-diabetic cases (225). The dataset is intended for reuse in method development (for example, machine-learning classifier training, feature-selection benchmarking, and oversampling/imputation research), for context-specific epidemiologic description and model validation in South Asian clinical settings, and as a teaching resource for reproducible biomedical-data workflows.
Thalassemia dataset covering clinical, socioeconomic, and mental health aspects
Abstract
Thalassemia is an inherited disorder of haemoglobin formation that requires lifelong blood transfusions and iron removal and brings both medical and emotional challenges. This dataset contains detailed information on 617 patients with thalassemia treated at a leading diagnostic centre in Pabna, Bangladesh. Demographic and socioeconomic data were gathered through face-to-face interviews, including age, gender, place of residence (urban or rural), education level, monthly household income, travel time to hospital, and how thalassemia affects school or work life. Clinical details, such as the type of thalassemia, age at diagnosis, whether the spleen was removed, transfusion frequency, the iron chelation plan, and how well patients adhered to it were taken from medical records. Laboratory results from the hospital’s central lab include pre transfusion haemoglobin, serum ferritin, red blood cell indices, reticulocyte count, and foetal haemoglobin percentage. Patients also rated their recent mental health as good or bad. Each record is assigned a unique code and stored in a single CSV file with twenty-three variables. The combination of medical, laboratory, social, and psychological information can support many uses, from building machine learning models to predict transfusion needs or therapy adherence to studying factors that drive health disparities and designing more patient centred care plans. • This dataset offers a unique, multidimensional view of 617 thalassemia patients by combining clinical, laboratory, demographic, socioeconomic, and mental health data into a single, structured resource. • This dataset’s mix of numerical, ordinal, and categorical features enables machine learning models to predict clinical outcomes like transfusion needs, iron chelation adherence, and mental health, supporting personalized treatment plans. • Combining self-reported mental health, school/work impact, and clinical measures like hemoglobin and transfusion frequency allows a comprehensive analysis of thalassemia’s physical and psychological burdens. • Researchers, healthcare professionals, and policymakers can use this dataset to explore the relationships between chronic illness management, psychological well-being, and social context, helping to improve patient-centered interventions and public health strategies.
A customized ensemble machine learning approach: predicting students' exam performance
Abstract
Accurately predicting students’ exam performance is crucial for fostering academic success and timely interventions. This study addresses the significant challenge of predicting whether a student will pass or fail based on key factors such as study hours and previous exam scores. Using a dataset of 500 students sourced from Kaggle, we introduce a novel customized ensemble machine learning model, combining Random Forest (RF) and AdaBoost classifiers with a custom-weighted soft voting method (weights of 0.2 for RF and 0.8 for AdaBoost). The model’s hyperparameters were optimized via GridSearchCV with 10-fold cross-validation, ensuring robustness. The performance of the ensemble model was evaluated using metrics like Cohen’s Kappa, achieving superior predictive accuracy compared to baseline models. Our findings indicate that the proposed model not only improves prediction accuracy but also reduces prediction time, offering practical implications for educators and policymakers to design tailored interventions for at-risk students, ultimately enhancing educational outcomes.
REDf: a deep learning model for short-term load forecasting to facilitate renewable integration and attaining the SDGs 7, 9, and 13
Abstract
Integrating renewable energy sources into the power grid is becoming increasingly important as the world moves towards a more sustainable energy future in line with the United Nations (UN) Sustainable Development Goal (SDG) 7 (Affordable and Clean Energy). However, the intermittent nature of renewable energy sources can make it challenging to manage the power grid and ensure a stable supply of electricity, which is crucial for achieving SDG 9 (Industry, Innovation and Infrastructure). In this article, we propose a deep learning model for predicting energy demand in a smart power grid, which can improve the integration of renewable energy sources by providing accurate predictions of energy demand. Our approach aligns with SDG 13 (Climate Action) on climate action, enabling more efficient management of renewable energy resources. We use long short-term memory networks, well-suited for time series data, to capture complex patterns and dependencies in energy demand data. The proposed approach is evaluated using four historical short-term energy demand data datasets from different energy distribution companies, including American Electric Power, Commonwealth Edison, Dayton Power and Light, and Pennsylvania-New Jersey-Maryland Interconnection. The proposed model is compared with three other state-of-the-art forecasting algorithms: Facebook Prophet, support vector regression, and random forest regression. The experimental results show that the proposed REDf model can accurately predict energy demand with a mean absolute error of 1.4%, indicating its potential to enhance the stability and efficiency of the power grid and contribute to achieving SDGs 7, 9, and 13. The proposed model also has the potential to manage the integration of renewable energy sources effectively.
Real-Time Detection of Forest Fires Using FireNet-CNN and Explainable AI Techniques
Abstract
This study presents FireNet-CNN, an advanced deep-learning model particularly designed for forest fire detection, which significantly surpasses existing methods in terms of reliability, efficiency, and interpretability. FireNet-CNN is compared to popular pre-trained models, including VGG16, VGG19, and Inception V3, across key performance metrics and consistently shows superior results, achieving 99.05% accuracy, 99.41% precision, and 98.28% recall. The model was evaluated using two augmented datasets: Dataset A and Dataset B, which consist of fire and non-fire images sourced from multiple video and image datasets. FireNet-CNN’s architecture, which includes 2.75 million parameters and a compact model size of 10.58 MB, has been meticulously optimized for fire detection tasks. As a consequence, the inference time of 0.95 seconds/image enables fast real-time deployment especially suitable for resource-constrained platforms like drones, remote sensors or other types of embedded systems in wooded regions. FireNet-CNN uses synthetic data augmentation based on Stable Diffusion to overcome the limitations of dataset size and class imbalance. This augmentation is critical as it helps the model accurately identify fire instances with a lower false positive rate, which is key for any real-time fire detection system where reliability and dependability are vital. To improve transparency and trust in safety-critical applications, FireNet-CNN incorporates the explainable AI (XAI) techniques, such as Grad-CAM and Saliency Maps. Despite encountering challenges such as reliance on synthetic data and issues of class imbalance, FireNet-CNN has demonstrated promising potential as a viable and effective solution for early wildfire detection. It offers significant insights for future research and practical applications in fire management and disaster response.
A Comprehensive dental dataset of six classes for deep learning based object detection study
Abstract
This article presents a dental dataset for the improvement of research on deep learning-based detection and classification of dental diseases. The dataset is consisted of 232 panoramic dental radiographs, categorized into six major classes: healthy teeth, caries, impacted teeth, infections, fractured teeth, and broken-down crowns/roots (BDC/BDR). The images were collected from three renowned private clinics in Dhaka, Bangladesh, with the help of an experienced dental practitioner who ensured the confidentiality of patients and high-quality data acquisition using a 64-megapixel Android phone camera. To enhance the value of the dataset for machine and deep learning applications, we applied Contrast-Limited Adaptive Histogram Equalization (CLAHE) for image enhancement and augmented the data. The images were annotated using the CVAT tool and reviewed by dental experts. This benchmark dataset is publicly available and provides a valuable resource for researchers in artificial intelligence, computer science, and dental informatics to promote interdisciplinary collaboration and the development of advanced algorithms for dental disease detection.
A novel integrated logistic regression model enhanced with recursive feature elimination and explainable artificial intelligence for dementia prediction
Abstract
Dementia is a major global health issue that significantly impacts millions of individuals, families, and societies worldwide, creating a substantial burden on healthcare systems. This study introduces a novel approach for predicting dementia by employing the Logistic Regression (LR) model, enhanced with Recursive Feature Elimination (RFE), applied to a unique dataset comprising 1000 patients, with 49.60% male and 50.40% female. The LR model, recognized for its simplicity and effectiveness in binary classification tasks, is optimized through RFE, a technique that iteratively eliminates less significant features to improve model performance. The model’s effectiveness was assessed using comprehensive metrics, including accuracy, precision, recall, F1-score, Matthews Correlation Coefficient (MCC), and Kappa score. Furthermore, SHapley Additive exPlanations (SHAP) values were employed to increase the interpretability of the model, providing insights into the most influential features for dementia prediction. To address the issue of overfitting, a standardization technique was implemented, which enhanced the model’s predictive performance. The findings of this study hold potential implications for early dementia detection, informing intervention strategies, and optimizing healthcare resource allocation. • Presents a dementia prediction model by utilizing the logistic regression enhanced with recursive feature elimination. • Enhance model accuracy and efficiency through recursive feature elimination. • Improve model interpretability with SHapley Additive exPlanations (SHAP) values. • Potential for early dementia prediction and intervention.
Explainable deep learning for diabetes diagnosis with DeepNetX2
Abstract
Diabetes is a leading health global health challenge because of its high blood sugar levels and the risk of extensive damage to other internal organs. Early and accurate identification of diabetes is important because it may cause other diseases including heart diseases and nerve damage. Despite the success of using machine learning, especially deep learning in automated diabetes diagnosis. These models are mostly black boxes which rarely offer comprehensive explanations and interpretations of the results. This study introduces DeepNetX2, a proposed custom deep neural network designed to overcome these challenges by integrating Explainable Artificial Intelligence (XAI) techniques, specifically Local Interpretable Model-agnostic Explanations (LIME) and Shapley Additive Explanations (SHAP). These techniques make the decision-making process of the model transparent, thereby increasing the credibility of the predictions. The proposed methodology entails a comprehensive data preprocessing technique that involves a customized Spearman’s correlation coefficient feature selection strategy. This preprocessing restricts complexity to only relevant features that promote effectiveness, instead of oversimplification, to the point of decreasing efficiency. DeepNetX2 was rigorously tested on three datasets: the PIMA dataset, local private dataset, and Type-2 diabetes dataset, achieving test accuracies of 94.81%, 97.87%, and 97.50%, respectively. These results demonstrate not only the superior performance of DeepNetX2 compared with existing models, but also its enhanced interpretability. Based on the proposed model, an appropriate and rapid strategy for diabetes prediction was developed which is very useful for improving diagnostic integrity and patient health.
A Medical Cyber-Physical System utilizing the Bayes algorithm for post-diagnosis patient supervision
Abstract
Among several basic human needs, medical treatment is one of the most prominent. However, since there need to be more doctors, nurses, and other medical facilities in many places, medical cyber-physical systems are quickly becoming a competitive alternative. One important use of these systems is for observation after a diagnosis. Instead of active observation by a caregiver, this can be easily done using various monitoring systems. However, most of the existing systems for this application are inflexible and need to consider the current challenges. On the other hand, these problems can be solved by intelligent and adaptive systems, which is now possible thanks to the growth of relevant technology, especially healthcare 4.0. Therefore, this article proposes an adaptive system based on the Bayes algorithm for performing medical interventions on patients, leading to a reduction in the dependence on caregivers, particularly in the post-diagnosis scenario.
A multimodal approach to cross-lingual sentiment analysis with ensemble of transformer and LLM
Abstract
Sentiment analysis is an essential task in natural language processing that involves identifying a text's polarity, whether it expresses positive, negative, or neutral sentiments. With the growth of social media and the Internet, sentiment analysis has become increasingly important in various fields, such as marketing, politics, and customer service. However, sentiment analysis becomes challenging when dealing with foreign languages, particularly without labelled data for training models. In this study, we propose an ensemble model of transformers and a large language model (LLM) that leverages sentiment analysis of foreign languages by translating them into a base language, English. We used four languages, Arabic, Chinese, French, and Italian, and translated them using two neural machine translation models: LibreTranslate and Google Translate. Sentences were then analyzed for sentiment using an ensemble of pre-trained sentiment analysis models: Twitter-Roberta-Base-Sentiment-Latest, bert-base-multilingual-uncased-sentiment, and GPT-3, which is an LLM from OpenAI. Our experimental results showed that the accuracy of sentiment analysis on translated sentences was over 86% using the proposed model, indicating that foreign language sentiment analysis is possible through translation to English, and the proposed ensemble model works better than the independent pre-trained models and LLM.
Priority based job scheduling technique that utilizes gaps to increase the efficiency of job distribution in cloud computing
Abstract
A growing number of services, accessible and usable by individuals and businesses on a pay-as-you-go basis, are being made available via cloud computing platforms. The business services paradigm in cloud computing encounters several quality of service (QoS) challenges, such as flow time, makespan time, reliability, and delay. To overcome these obstacles, we first designed a resource management framework for cloud computing systems. This framework elucidates the methodology of resource management in the context of cloud job scheduling. Then, we study the impact of a Virtual Machine's (VM's) physical resources on the consistency with which cloud services are executed. After that, we developed a priority-based fair scheduling (PBFS) algorithm to schedule jobs so that they have access to the required resources at optimal times. The algorithm has been devised utilizing three key characteristics, namely CPU time, arrival time, and job length. For optimal scheduling of cloud jobs, we also devised a backfilling technique called Earliest Gap Shortest Job First (EG-SJF), which prioritizes filling in schedule gaps in a specific order. The simulation was carried out with the help of the CloudSim framework. Finally, we compare our proposed PBFS algorithm to LJF, FCFS, and MAX–MIN and find that it achieves better results in terms of overall delay, makespan time, and flow time.
Evaluating the Performance of a Visual Support System for Driving Assistance using a Deep Learning Algorithm
Abstract
The issue of road accidents endangering human life has become a global concern due to the rise in traffic volumes. This article presents the evaluation of an object detection model for University of Malaysia Pahang (UMP) roadside conditions, focusing on the detection of vehicles, motorcycles, and traffic lamps. The dataset consists of the driving distance from Hospital Pekan to the University of Malaysia Pahang. Around one thousand images were selected in Roboflow for the train dataset. The model utilises the YOLO V8 deep learning algorithm in the Google Colab environment and is trained using a custom dataset managed by the Roboflow dataset manager. The dataset comprises a diverse set of training and validation images, capturing the unique characteristics of Malaysian roads. The train model's performance was assessed using the F1 score, precision, and recall, with results of 71%, 88.2%, and 84%, respectively. A comprehensive comparison with validation results has shown the efficacy of the proposed model in accurately detecting vehicles, motorcycles, and traffic lamps in real-world Malaysian road scenarios. This study contributes to the improvement of intelligent transportation systems and road safety in Malaysia.
An automated materials and processes identification tool for material informatics using deep learning approach
Abstract
This article reports a tool that enables Materials Informatics, termed as MatRec, via a deep learning approach. The tool captures data, makes appropriate domain suggestions, extracts various entities such as materials and processes, and helps to establish entity-value relationships. This tool uses keyword extraction, a document similarity index to suggest relevant documents, and a deep learning approach employing Bi-LSTM for entity extraction. For example, materials and processes for electrical charge storage under an electric double layer capacitor (EDLC) mechanism are demonstrated herewith. A knowledge graph approach finds and visualizes different latent knowledge sets from the processed information. The MatRec received an F1 score of 9̃6% for entity extraction, 8̃3% for material-value relationship extraction, and 8̃7% for process-value relationship extraction, respectively. The proposed MatRec could be extended to solve material selection issues for various applications and could be an excellent tool for academia and industry.
Aspect-based Sentiment Analysis Model for Evaluating Teachers' Performance from Students' Feedback
Abstract
Evaluating teachers' performance is a fundamental pillar of educational enhancement, guiding the evolution of pedagogical practices and fostering enriched learning environments. This study pioneers an innovative approach by harnessing sentiment analysis within an aspect-based framework to decipher the intricate emotional nuances embedded within students' feedback. By categorizing sentiments as positive, negative, and neutral, we delve into the diverse perceptions of teaching aspects, offering a multifaceted portrait of educators' contributions. Through meticulous data collection, preprocessing, and a deep learning sentiment analysis model, we dissected student comments into distinct teaching aspects. The subsequent sentiment analysis unearthed positive, negative, and neutral sentiments. Positive sentiments highlighted strengths and effective communication, while negative sentiments illuminated areas for growth. Neutral sentiments provided contextual equilibrium, forming a holistic tapestry of teachers' performance. The proposed model achieved 86\% F1 score for classifying sentiments into three classes.
A comprehensive dataset for aspect-based sentiment analysis in evaluating teacher performance
Abstract
Teacher performance evaluation is an essential task in the field of education. In recent years, aspect-based sentiment analysis (ABSA) has emerged as a promising technique for evaluating teaching performance by providing a more nuanced analysis of student evaluations. This article presents a novel approach for creating a large-scale dataset for ABSA of teacher performance evaluation. The dataset was constructed by collecting student feedback from American International University-Bangladesh and then labeled by undergraduate-level students into three sentiment classes: positive, negative, and neutral. The dataset was carefully cleaned and preprocessed to ensure data quality and consistency. The final dataset contains over 2,000,000 student feedback instances related to teacher performance, making it one of the largest datasets for ABSA of teacher performance evaluation. This dataset can be used to develop and evaluate ABSA models for teacher performance evaluation, ultimately leading to better feedback and improvement for educators. The results of this study demonstrate the usefulness and effectiveness of ABSA in evaluating teacher performance and highlight the importance of creating high-quality datasets for this task.
Yus - A Deep Learning Algorithm for Collision Avoidance through Object and Vehicle Detection
Abstract
One of the safety features that can alert drivers to the presence of other vehicles and reduce the risk of collisions is vehicle detection. In this study, the objective was to create a system for detecting vehicles, motorcycles, and traffic signals on the roads in University Malaysia Pahang using object detection techniques. The video was taken through Go-Pro camera to capture video footage of traffic objects on the roads in the district, which was then analysed using the YOLO-V8 deep learning algorithm. The system was trained on a pre-existing dataset of 1,068 images, with 70% of the dataset used for training and 30% for testing. After conducting a performance validation, the system achieved a mean average precision of 88.2% on training dataset and was able to detect different types of vehicles such as cars, motorcycles, and traffic lights. The results of this study could be beneficial for road safety authorities and researchers interested in developing intelligent transportation systems.
Distributed Ledger Technology Based Integrated Healthcare Solution for Bangladesh
Abstract
Healthcare data is highly sensitive and must be safeguarded. Personal and sensitive data, such as names and addresses, is stored in Encrypted Electronic Health Records (EHRs). This paper proposes a Blockchain-based distributed application platform for Bangladesh’s public and private healthcare service providers. The proposed application framework enables users to create secure digital agreements for commerce or collaboration by leveraging data immutability and smart contracts. As a result, all stakeholders can collaborate securely over the same Blockchain network, taking advantage of their data’s openness and read/write nature. The proposed application is made up of various application interfaces for various stakeholders. The proposed solution employs Hyperledger Fabric and Blockchain to ensure data integrity, privacy, permissions, and service availability. In the application portal, each user has a profile. The creation of a unique identity for each user, as well as the establishment of digital information centers across the country, has greatly aided the process. This application collects health data from each user in a systematic manner, which is useful for research institutes and healthcare-related organizations. For this application, a national data warehouse in Bangladesh is feasible, and various healthcare-related analyses can be performed using the collected data, assisting the strategy and planning department in making informed decisions regarding the healthcare sector in Bangladesh. Because Bangladesh has both public and private healthcare providers, a simple digital strategy is essential for all organizations to accomplish their services. This study proposes a solution to achieve this goal.
4D: A Real-Time Driver Drowsiness Detector Using Deep Learning
Abstract
There are a variety of potential uses for the classification of eye conditions, including tiredness detection, psychological condition evaluation, etc. Because of its significance, many studies utilizing typical neural network algorithms have already been published in the literature, with good results. Convolutional neural networks (CNNs) are employed in real-time applications to achieve two goals: high accuracy and speed. However, identifying drowsiness at an early stage significantly improves the chances of being saved from accidents. Drowsiness detection can be automated by using the potential of artificial intelligence (AI), which allows us to assess more cases in less time and with a lower cost. With the help of modern deep learning (DL) and digital image processing (DIP) techniques, in this paper, we suggest a CNN model for eye state categorization, and we tested it on three CNN models (VGG16, VGG19, and 4D). A novel CNN model named the 4D model was designed to detect drowsiness based on eye state. The MRL Eye dataset was used to train the model. When trained with training samples from the same dataset, the 4D model performed very well (around 97.53% accuracy for predicting the eye state in the test dataset). The 4D model outperformed the performance of two other pretrained models (VGG16, VGG19). This paper explains how to create a complete drowsiness detection system that predicts the state of a driver’s eyes to further determine the driver’s drowsy state and alerts the driver before any severe threats to road safety.
ReSTiNet: An Efficient Deep Learning Approach to Improve Human Detection Accuracy
Abstract
Human detection is an important task in computer vision. It is one of the most important tasks in global security and safety monitoring. In recent days, Deep Learning has improved human detection technology. Despite modern techniques, there are very few optimal techniques to construct networks with a small size, deep architecture, and fast training time while maintaining accuracy. ReSTiNet is a novel small convolutional neural network that overcomes the problems of network size, detection speed, and accuracy. The developed ReSTiNet contains fire modules by evaluating their number and position in the network to minimize the model parameters and network size. To improve the detection speed and accuracy of ReSTiNet, the residual block within the fire modules is carefully designed to increase the feature propagation and maximize the information flow in the network. The developed approach compresses the well-known Tiny-YOLO architecture while improving the following features: (i) small model size, (ii) faster detection speed, (iii) resolution of overfitting, and (iv) better performance than other compact networks such as SqueezeNet and MobileNet in terms of mAP on the Pascal VOC and MS COCO datasets. ReSTiNet is 10.7 MB, five times smaller than Tiny-YOLO. On Tesla k80, mAP is 27.3% for MS COCO and 63.74% for PASCAL VOC. The validation of the proposed ReSTiNet model has been done on INRIA person dataset using the Tesla K80. •All the necessary steps, algorithms, and mathematical formulas for building the net- work are provided. •The network is small in size but has a faster detection speed with high accuracy.
ReSTiNet: On Improving the Performance of Tiny-YOLO-Based CNN Architecture for Applications in Human Detection
Abstract
Human detection is a special application of object recognition and is considered one of the greatest challenges in computer vision. It is the starting point of a number of applications, including public safety and security surveillance around the world. Human detection technologies have advanced significantly in recent years due to the rapid development of deep learning techniques. Despite recent advances, we still need to adopt the best network-design practices that enable compact sizes, deep designs, and fast training times while maintaining high accuracies. In this article, we propose ReSTiNet, a novel compressed convolutional neural network that addresses the issues of size, detection speed, and accuracy. Following SqueezeNet, ReSTiNet adopts the fire modules by examining the number of fire modules and their placement within the model to reduce the number of parameters and thus the model size. The residual connections within the fire modules in ReSTiNet are interpolated and finely constructed to improve feature propagation and ensure the largest possible information flow in the model, with the goal of further improving the proposed ReSTiNet in terms of detection speed and accuracy. The proposed algorithm downsizes the previously popular Tiny-YOLO model and improves the following features: (1) faster detection speed; (2) compact model size; (3) solving the overfitting problems; and (4) superior performance than other lightweight models such as MobileNet and SqueezeNet in terms of mAP. The proposed model was trained and tested using MS COCO and Pascal VOC datasets. The resulting ReSTiNet model is 10.7 MB in size (almost five times smaller than Tiny-YOLO), but it achieves an mAP of 63.74% on PASCAL VOC and 27.3% on MS COCO datasets using Tesla k80 GPU.
Evaluating keyphrase extraction algorithms for finding similar news articles using lexical similarity calculation and semantic relatedness measurement by word embedding
Abstract
A textual data processing task that involves the automatic extraction of relevant and salient keyphrases from a document that expresses all the important concepts of the document is called keyphrase extraction. Due to technological advancements, the amount of textual information on the Internet is rapidly increasing as a lot of textual information is processed online in various domains such as offices, news portals, or for research purposes. Given the exponential increase of news articles on the Internet, manually searching for similar news articles by reading the entire news content that matches the user's interests has become a time-consuming and tedious task. Therefore, automatically finding similar news articles can be a significant task in text processing. In this context, keyphrase extraction algorithms can extract information from news articles. However, selecting the most appropriate algorithm is also a problem. Therefore, this study analyzes various supervised and unsupervised keyphrase extraction algorithms, namely KEA, KP-Miner, YAKE, MultipartiteRank, TopicRank, and TeKET, which are used to extract keyphrases from news articles. The extracted keyphrases are used to compute lexical and semantic similarity to find similar news articles. The lexical similarity is calculated using the Cosine and Jaccard similarity techniques. In addition, semantic similarity is calculated using a word embedding technique called Word2Vec in combination with the Cosine similarity measure. The experimental results show that the KP-Miner keyphrase extraction algorithm, together with the Cosine similarity calculation using Word2Vec (Cosine-Word2Vec), outperforms the other combinations of keyphrase extraction algorithms and similarity calculation techniques to find similar news articles. The similar articles identified using KPMiner and the Cosine similarity measure with Word2Vec appear to be relevant to a particular news article and thus show satisfactory performance with a Normalized Discounted Cumulative Gain (NDCG) value of 0.97. This study proposes a method for finding similar news articles that can be used in conjunction with other methods already in use.
A machine learning approach for Bengali handwritten vowel character recognition
Abstract
Recognition of handwritten characters is complex because of the different shapes and numbers of characters. Many handwritten character recognition strategies have been proposed for both English and other major dialects. Bengali is generally considered the fifth most spoken local language in the world. It is the official and most widely spoken language of Bangladesh and the second most widely spoken among the 22 posted dialects of India. To improve the recognition of handwritten Bengali characters, we developed a different approach in this study using face mapping. It is quite effective in distinguishing different characters. The real highlight is that the recognition results are more efficient than expected with a simple machine learning technique. The proposed method uses the Python library Scikit-Learn, including NumPy, Pandas, Matplotlib, and support vector machine (SVM) classifier. The proposed model uses a dataset derived from the BanglaLekha isolated dataset for the training and testing part. The new approach shows positive results and looks promising. It showed accuracy up to 94% for a particular character and 91% on average for all characters.
Application of machine learning algorithms to predict the thyroid disease risk: an experimental comparative study
Abstract
Thyroid disease is the general concept for a medical problem that prevents one's thyroid from producing enough hormones. Thyroid disease can affect everyone-men, women, children, adolescents, and the elderly. Thyroid disorders are detected by blood tests, which are notoriously difficult to interpret due to the enormous amount of data necessary to forecast results. For this reason, this study compares eleven machine learning algorithms to determine which one produces the best accuracy for predicting thyroid risk accurately. This study utilizes the Sick-euthyroid dataset, acquired from the University of California, Irvine's machine learning repository, for this purpose. Since the target variable classes in this dataset are mostly one, the accuracy score does not accurately indicate the prediction outcome. Thus, the evaluation metric contains accuracy and recall ratings. Additionally, the F1-score produces a single value that balances the precision and recall when an uneven distribution class exists. Finally, the F1-score is utilized to evaluate the performance of the employed machine learning algorithms as it is one of the most effective output measurements for unbalanced classification problems. The experiment shows that the ANN Classifier with an F1-score of 0.957 outperforms the other nine algorithms in terms of accuracy.
Predicting Young Imposter Syndrome Using Ensemble Learning
Abstract
Background . Imposter syndrome (IS), associated with self‐doubt and fear despite clear accomplishments and competencies, is frequently detected in medical students and has a negative impact on their well‐being. This study aimed to predict the students’ IS using the machine learning ensemble approach. Methods . This study was a cross‐sectional design among medical students in Bangladesh. Data were collected from February to July 2020 through snowball sampling technique across medical colleges in Bangladesh. In this study, we employed three different machine learning techniques such as neural network, random forest, and ensemble learning to compare the accuracy of prediction of the IS. Results . In total, 500 students completed the questionnaire. We used the YIS scale to determine the presence of IS among medical students. The ensemble model has the highest accuracy of this predictive model, with 96.4%, while the individual accuracy of random forest and neural network is 93.5% and 96.3%, respectively. We used different performance matrices to compare the results of the models. Finally, we compared feature importance scores between neural network and random forest model. The top feature of the neural network model is Y7, and the top feature of the random forest model is Y2, which is second among the top features of the neural network model. Conclusions . Imposter syndrome is an emerging mental illness in Bangladesh and requires the immediate attention of researchers. For instance, in order to reduce the impact of IS, identifying key factors responsible for IS is an important step. Machine learning methods can be employed to identify the potential sources responsible for IS. Similarly, determining how each factor contributes to the IS condition among medical students could be a potential future direction.
Sentence Boundary Extraction from Scientific Literature of Electric Double Layer Capacitor Domain: Tools and Techniques
Abstract
Given the growth of scientific literature on the web, particularly material science, acquiring data precisely from the literature has become more significant. Material information systems, or chemical information systems, play an essential role in discovering data, materials, or synthesis processes using the existing scientific literature. Processing and understanding the natural language of scientific literature is the backbone of these systems, which depend heavily on appropriate textual content. Appropriate textual content means a complete, meaningful sentence from a large chunk of textual content. The process of detecting the beginning and end of a sentence and extracting them as correct sentences is called sentence boundary extraction. The accurate extraction of sentence boundaries from PDF documents is essential for readability and natural language processing. Therefore, this study provides a comparative analysis of different tools for extracting PDF documents into text, which are available as Python libraries or packages and are widely used by the research community. The main objective is to find the most suitable technique among the available techniques that can correctly extract sentences from PDF files as text. The performance of the used techniques Pypdf2, Pdfminer.six, Pymupdf, Pdftotext, Tika, and Grobid is presented in terms of precision, recall, f-1 score, run time, and memory consumption. NLTK, Spacy, and Gensim Natural Language Processing (NLP) tools are used to identify sentence boundaries. Of all the techniques studied, the Grobid PDF extraction package using the NLP tool Spacy achieved the highest f-1 score of 93% and consumed the least amount of memory at 46.13 MegaBytes.
Recommending Research Articles: A Multi-Level Chronological Learning-Based Approach using Unsupervised Keyphrase Extraction and Lexical Similarity Calculation
Abstract
A research article recommendation approach aims to recommend appropriate research articles to analogous researchers to help them better grasp a new topic in a particular research area. Due to the accessibility of research articles on the web, it is tedious to recommend a relevant article to a researcher who strives to understand a particular article. Most of the existing approaches for recommending research articles are metadata-based, citation-based, bibliographic coupling-based, content-based, and collaborative filtering-based. They require a large amount of data and do not recommend reference articles to the researcher who wants to understand a particular article going through the reference articles of that particular article. Therefore, an approach that can recommend reference articles for a given article is needed. In this paper, a new multi-level chronological learning-based approach is proposed for recommending research articles to understand the topics/concepts of an article in detail. The proposed method utilizes the TeKET keyphrase extraction technique, among other unsupervised techniques, which performs better in extracting keyphrases from the articles. Cosine and Jaccard similarity measures are employed to calculate the similarity between the parent article and its reference articles using the extracted keyphrases. The cosine similarity measure outperforms the Jaccard similarity measure for finding and recommending relevant articles to understand a particular article. The performance of the recommendation approach seems satisfactory, with an NDCG value of 0.87. The proposed approach can play an essential role alongside other existing approaches to recommend research articles.
Study of Keyword Extraction Techniques for Electric Double-Layer Capacitor Domain Using Text Similarity Indexes: An Experimental Analysis
Abstract
Keywords perform a significant role in selecting various topic‐related documents quite easily. Topics or keywords assigned by humans or experts provide accurate information. However, this practice is quite expensive in terms of resources and time management. Hence, it is more satisfying to utilize automated keyword extraction techniques. Nevertheless, before beginning the automated process, it is necessary to check and confirm how similar expert‐provided and algorithm‐generated keywords are. This paper presents an experimental analysis of similarity scores of keywords generated by different supervised and unsupervised automated keyword extraction algorithms with expert‐provided keywords from the electric double layer capacitor (EDLC) domain. The paper also analyses which texts provide better keywords such as positive sentences or all sentences of the document. From the unsupervised algorithms, YAKE, TopicRank, MultipartiteRank, and KPMiner are employed for keyword extraction. From the supervised algorithms, KEA and WINGNUS are employed for keyword extraction. To assess the similarity of the extracted keywords with expert‐provided keywords, Jaccard, Cosine, and Cosine with word vector similarity indexes are employed in this study. The experiment shows that the MultipartiteRank keyword extraction technique measured with cosine with word vector similarity index produces the best result with 92% similarity with expert‐provided keywords. This study can help the NLP researchers working with the EDLC domain or recommender systems to select more suitable keyword extraction and similarity index calculation techniques.
IoT (Internet of Things) - Based Smart Garbage Management System: A Proposal for major Cities of Bangladesh
Abstract
IoT –internet of things has become a buzzword nowadays. There are many IoT based researches but researches on garbage management system based on IoT are not sufficient. Insufficient and inefficient garbage management system causes severe environmental problem. It also makes the air toxic. This problem has become a common problem in the world especially in Bangladesh. Dhaka city, the capital of Bangladesh lacks well organized and efficient garbage management system. Maximum roads of Dhaka city are surrounded by garbage. The bad smell of garbage affects people’s mental health, inhaling toxic causes many diseases. Lack of dustbins, throwing of garbage here and there, misuse of dustbins are making city life very unhealthy and also causes a threat to environment. The dustbins are being stolen or damaged which is also a great problem. In this paper, we proposed about an efficient garbage management system based on IoT. This research works aims to provide a minimal solution to this problem using the IoT technology. We propose for a smart garbage system, which consists of sensors, RFID, IR sensors, admin and user website, Wi-Fi module etc. These smart bins will monitor the level of garbage when it will reach 75% of its capacity, it will give notification to the admin website, so the authority concerned can collect the garbage from the bins timely and there will be no overflow of garbage as the authority will get notified earlier. There will be a feature in user website that will let the user know about the nearest smart garbage bins current condition, so if there is any condition that the garbage bin of their place is full they can use the nearest bin. This research work also aims to have secured smart garbage bins, as there is chance the bins to be stolen and damaged so in this research we talk about security of the sensors and the bins will have cement body. So this research is for implementing an efficient garbage management system which will reduce expense on this sector, misuse of bins. Making a clean country, pollution free environment with an efficient and well organized garbage management system can bring a new era. It is expected that the proposed garbage management system will reduce financial cost in this sector as well as reduce problems related to waste management.
Demography of Startup Software Companies: An Empirical Investigation on the Success and Failure
Abstract
Startup software company can be envisage as a company committed to deliver innovative solution to market with four primary objectives, namely, product, service, process and platform.Software startups are quite distinct from traditional mature software companies.These companies develop state-of-the-art, software-intensive products within limited time frames and with few resources.They also introduce new challenges relevant to software engineering research which in turn lead to ample set of empirical investigation on the topic.This research is targeted to offer a systematic literature survey on the scientific investigations carried out on Startup software companies, specifically concentrating on their success and failure parameters and strategies to encounter symptoms of failure.
Intelligent Tour Planning System using Crowd Sourced Data
Abstract
To observe the beauty of nature and to visit various places around the world, a vast number of tourists visit different countries and many tourist attraction sites now-a-days.But Most of the tourist places have failed to introduce itself as a tourist destination to the visitor due to lack of proper information and proper guideline to visit there.This paper tries to focus on some problems in the tourism industry and try to solve those problems using crowd sourced data with some customized algorithms.Some of the main problems are the lack of information about a destination tourist spot, combination on budget to visit the spot, time of travels etc.We proposed a customize algorithm which will provide maximum suggestion to visit a place with nearest all sub place based on user destination within their given budget and time.Using our method, user can choose the most suitable plan for them to visit those places.
Simple Approach to Traffic Update System
Abstract
Now a day's traffic jam is a very common problem in all the cities.Every day people spend valuable times uselessly because of traffic jam.We have proposed a system that came up with a solution which can reduce traffic jam dramatically.This is a web based application.People will just check the status of the route they are heading for whether the route consisting medium, heavy or less traffic.Moreover people can check all the alternative routes for the current traffic status that can help them taking decision for a traffic jam free route.
Face Recognition using Eigenfaces
Abstract
We tried to develop a real time face detection and recognition system which uses an "appearance-based" approach.For detection purpose we used Viola Jones algorithm.To recognize face we worked with Eigen Faces which is a PCA based algorithm.In a real time to recognize a face we need a data training set.For data training set we took five images of each person and manipulated the Eigen values to match the known individual.
Conference Papers
An Advanced Multi-Input LSTM Framework with Attention for Predicting the Risk Level of Cardiovascular Disease
Abstract
Cardiovascular disease (CVD) continues to be the leading cause of mortality globally. There is a need for accurate and clinically interpretable predictive systems for CVD. In this paper, we propose a multi-input Long Short-Term Memory (LSTM) model with an attention mechanism for predicting CVD, enhanced with uncertainty quantification via Monte Carlo Dropout and Bayesian-inspired techniques. To bridge predictive modeling with patient care, we further introduce a digital twin simulation for patient trajectory forecasting. The system integrates explainability tools, including attention heatmaps, SHAP, and LIME, alongside calibration analysis through reliability diagrams and Expected Calibration Error (ECE). Experimental results demonstrate strong predictive performance (AUC 0.77–0.81), reliable uncertainty estimates, and interpretable outputs, supporting its potential for clinical decision support.
DysLSTM: A Deep Sequential Framework for Dyslexia Prediction with Comparative Analysis of Baseline and Advanced Models
Abstract
Dyslexia is one of the most common types of learning disorders, affecting both the reading and decoding of words in children and adults. The early detection is of paramount importance for timely intervention and yet the widely-used detection approaches in traditional systems are expensive, time-consuming and subjective. In this work, we propose DysLSTM: a robust deep learning framework that relies on Long Short-Term Memory (LSTM) networks to accurately predict dyslexia. We trained and tested it on an open dataset, collecting 3,644 individuals, with elaborate data pre-processing, feature mapping, scaling and balancing. Experimental results on multiple baseline machine learning (Logistic Regression, Random Forest, SVM) and deep learning (CNN, GRU, RNN) models are shown that DysLSTM outperforms all of them with 97% accuracy and high precision-recall scores as well as a Cohen's Kappa score of 0.91, which suggests that DysLSTM has nearly perfect agreement.
CAFC: A Convolutional Neural Approach for Sentiment Recognition in Speech Signals
Abstract
Sentiment analysis using speech is an emerging field that leverages vocal cues such as pitch, tone, and rhythm to capture human emotions beyond textual content. This paper shows audio-based sentiment classification using Mel-Frequency Cepstral Coefficients (MFCCs) and compares the performance of traditional machine learning (ML) and deep learning (DL) models introduced in this study. A novel Convolutional Acoustic-Feature Classifier (CAFC) is proposed to address the limitations of existing approaches which is based on a 1D Convolutional Neural Network (CNN) architecture optimized for sequential acoustic features. The proposed CAFC model can automatically capture temporal and spectral dependence relationships in speech, and it is more efficient and accurate compared with current supervised methods in audio sentiment analysis. The CAFC model was validated using 5-fold cross-validation which demonstrated that CAFC not only consistently outperforms all baseline models but also achieving an accuracy of 96%, with superior precision, recall, and F1-score. For sentiment recognition in speech the findings indicated the effectiveness of CNN-based architectures and underscore the importance of leveraging acoustic features for robust human–computer interaction systems.
Hybrid CNN-BiLSTM Model for Accurate Forecasting of Atmospheric Ozone Concentrations: Enhancing Temporal and Spatial Dependencies in Environmental Predictions
Abstract
Forecasting ozone concentrations in the atmosphere is important for air quality management and public health protection, especially in urban areas. The objective of this research is to propose a hybrid CNN-BiLSTM model that simultaneously captures long-term spatio-temporal dependencies in time series as well as space-time, multi-channel feature representations for realtime high-dimensional air quality monitoring data e.g. for ozone, particulate matter, and fine particulate matter. Convolutional neural network (CNN) layers are used to extract automatic highlevel spatial features from multiple input features, while BiLSTMs provide both forward and backward temporal properties to facilitate sequence prediction and memory-based reasoning. The preprocessing, encoding, model training, and evaluation methodology use standard indices of performance, including mean squared error (MSE), mean absolute error (MAE), root mean squared error (RMSE), and coefficient of determination (R2). Experimental results demonstrate that the proposed CNN-BiLSTM model consistently outperforms individual CNN, LSTM, and BiLSTM models in terms of prediction accuracy and generalization capability. The results provide clear and important indications of the benefits associated with the introduction of spatial-temporal cognition into the unified framework. Beyond improved ozone concentration forecasting, this study contributes a scalable and transferable deep learning approach applicable to broader time-series prediction tasks in environmental monitoring and real-time decision-making systems.
Hyperparameter-Optimized Stacking with SMOTETomek Balancing for Robust Multi-Cancer Risk Prediction
Abstract
Heterogeneous lifestyle, genetic, and environmental factors are difficult to predict risk of cancer because of the interplay of features and extreme imbalance in classes of cancer. This paper is a proposal of a two-phase optimized stacking model to multi-cancer risk prediction, specifically focusing on Lung, Breast, Colon, Prostate and Skin cancers using tabular population level data. The analysis of 2,000 people with 21 risk related attributes was performed in a structured preprocessing pipeline which involved feature selection, categorical encoding, stratified data splitting, and imbalance correction with the SMOTETomek resampling. During the initial stage, systematic benchmarking of multiple baseline machine learning models was done. During the second phase, hyperparameter optimized most competitive learners were stacked to form an ensemble, and a meta-learner was used to combine the predictions of hyperparameter optimized base models to use their complementary advantages. The proposed stacking method was always better than single baseline models, with better robustness and cross-classification of cancer. The reliability of the model was also assessed through 5 fold cross-validation and the predictive behavior across folds was found to be stable. This model achieved substantial agreement beyond chance with Cohen's kappa. The findings suggest that the offered framework offers a high quality and computationally efficient model of multi-cancer risk assessment, and it has high chances to be implemented in resource restricted medical conditions to aid early screening and preventive decision making.
Optimized Ensemble Learning with SMOTE for Gastric Cancer Classification
Abstract
Gastric cancer (GC) is a common and mortal cancer worldwide, and the classification of its subtypes can be used as independent factors in determining the prognosis and therapeutic strategies. However, due to overlapping morphological characteristics as well as class imbalance in clinical datasets, it is difficult to distinguish the histological subtypes accurately. The objective of this study is to develop an optimal ensemble learning framework that balances the data through SMOTE and applies a hybrid soft voting ensemble that is formed with Random Forest, XGBoost and Logistic Regression classifiers. This study compare the results between the baseline and ensemble models. The findings indicate a noteworthy enhancement produced by the use of class rebalancing for minority classes. The ensemble model also performed better than all other single models as indicated by an accuracy of 90%. These results demonstrate that ensemble learning integrated with class balancing approach can achieve an accurate and equitable classification of gastric cancer subtypes, contributing to guiding precise and personalized clinical decisions.
AI in Education: PCA-IG and VAE-Powered Insights Into Chatgpt's Impact on Learners
Abstract
Artificial Intelligence (AI) technologies are transforming the way students learn and interact with classroom activities. However, there is a lack of research on student perceptions and performances, particularly among developing nations such as Bangladesh. In this research, the study conducts a survey among 3,512 students to evaluate AI technologies (such as ChatGPT) within the educational system. To extract meaningful insights, we propose a robust Machine Learning (ML) framework that combines Principal Component Analysis with Information Gain (PCA-IG) for feature selection and a Variational Autoencoder (VAE) for oversampling imbalanced class distributions. Our findings reveal that factors such as support for critical thinking and data privacy concerns significantly influence student's perceived learning outcomes. Among various classifiers tested, the K-Nearest Neighbors (KNN) model enhanced with VAEbased oversampling achieved the highest accuracy of 92.54 % and strong AUC scores across all classes, outperforming traditional approaches. These results demonstrate ChatGPT's revolutionary potential in higher education when applied sensibly and with the proper privacy and ethical safeguards. For educators and policymakers wishing to successfully incorporate AI tools into educational settings, the proposed technique provides a scalable approach to examine AI-driven learning environments.
Lightweight, Stain-Normalized White Blood Cell Classification with Explainability and Web-Deployed Prototype
Abstract
Accurate white blood cell (WBC) classification from stained microscopy images is critical for hematological diagnosis but is impeded by stain variability and class imbalance. This work aims to deliver a lightweight, stain-robust, and explainable WBC classifier with a publicly available demo to support reproducible evaluation. We normalize input images using Vahadane stain normalization to reduce color heterogeneity and apply targeted data augmentation to address class imbalance. An EfficientV2-B1 backbone (≈ 7 million parameters) is trained on the augmented dataset, and model decisions are interpreted using Grad-CAM heatmaps. Statistical robustness is quantified through bootstrap resampling with 1,500 samples to generate 95% confidence intervals for accuracy. The resulting model attains 98.69% test accuracy, with bootstrap analysis indicating a 95% confidence interval of 98.22-99.14%, and Grad-CAM visualizations corroborate that predictions rely on biologically plausible regions. We deploy an interactive web application that displays the original image, the Vahadane-normalized image, Grad-CAM overlays, and per-class confidence scores to promote transparency and practical use. Overall, the proposed pipeline combines stain normalization, augmentation, an efficient architecture, rigorous validation, and explainability to produce a robust, highperforming WBC classification system suitable for research and clinical translation.
Hybrid Framework for Tomato Leaf Disease Detection Employing Densenet121 and Ensemble Learning Techniques
Abstract
Tomato leaf diseases significantly reduce yield and quality in commercial production, making early and reliable diagnosis essential. This paper presents a hybrid tomato leaf disease classification framework that combines deep feature extraction using a pre-trained DenseNet121 model with multiple traditional machine learning classifiers and soft-voting ensemble. The dataset consists of 622 tomato leaf images from six classes (five disease categories and one healthy class), which are augmented through rotations, flips, brightness adjustment, and random zooming to address class imbalance and increase variability. Each image is resized to$224 \times 224$pixels and normalized before being passed through DenseNet121 to obtain high-level feature representations. These features are then used to train five classifiers: k-nearest neighbors, support vector machine, multilayer perceptron, LightGBM, and random forest. Their probabilistic outputs are fused using a soft-voting strategy to obtain the final prediction. Experimental results show that the ensemble model achieves 99.49% accuracy, outperforming all individual classifiers and several recent state-of-the-art methods on the same task. The proposed framework demonstrates that combining deep feature extraction with lightweight traditional classifiers provides an accurate and computationally efficient solution for practical tomato leaf disease detection and decision support in precision agriculture.
Exploring Mental Health Through Sentiment Analysis: Predictive Models Using Natural Language Processing (NLP)
Abstract
Mental health has turned out to be a global issue with the presence of digital communication, and there is an increasing scope for textual data to be mined on one’s emotional and physiological health. This research examines the combination of sentiment analysis and a predictive model using NLP to detect and analyze mental health conditions from text data. The primary objective of this research is to evaluate predictive models that would be capable of detecting mental health conditions from sentiment analysis of text data. Involving in this task in the testing and comparing different types of NLP architectures including Transformer models and traditional ML models to see which of these methods is better in identifying signs of mental health conditions. This study uses a rich dataset that is collected from multiple sources like social media text, personal diaries, etc. This study trained multiple ML algorithms such as the Transformer model (RoBerTa, BERT, and DeBERTa) and Non-Transformer model (Naive Bayes, Random Forest, Logistic Regression) to determine the mental health classification of the text data. This work evaluated multiple ML algorithms for detecting mental health categories, highlighting the one best model that consistently best performs with high \(F_1\)-scores among the rest. The model accuracy was high, and errors were few. As a whole, this study makes an important contribution to mental health and offers the basis for developing tools that could help both professionals and individuals.
Data-Driven Diagnosis of Monkeypox: Insights from Machine Learning Models
Abstract
In this study, machine learning models were implemented, and their efficiency was evaluated to predict monkeypox diagnosis based on patient symptom data. The primary aim was to compare the performance of various classifiers, including Support Vector Machine (SVM), Decision Tree, Random Forest, Logistic Regression, XGBoost, and LightGBM. Initially, SVM demonstrated a high single accuracy score of 70.16%, but cross-validation revealed a significant drop to an average accuracy of 63.64%, suggesting poor generalization. The feature importance analysis from tree-based models was used to identify and drop less relevant features, but the improvements in performance were marginal, indicating that no single symptom strongly predicted the diagnosis. An important preprocessing step involved removing rows where both Systemic Illness was absent and monkeypox was negative, which reduced noise in the dataset and significantly improved model performance by 4–8%, particularly for Logistic Regression which has demonstrated an improvement of 8%. Gradient-boosting algorithms, such as LightGBM and XGBoost, performed closely to Logistic Regression. This research emphasizes the importance of effective data cleaning, model selection, and the use of advanced algorithms for handling complex medical datasets. Future work could explore more sophisticated feature engineering and deep learning techniques to enhance prediction accuracy in monkeypox diagnostics.
Machine Learning Models for Reliable Thyroid Cancer Recurrence Prediction: A Comparative Analysis
Abstract
Among all endocrine malignancies, thyroid cancer is the common endocrine malignant in our body. Thyroid cancer can be developed from two types of cells where the first type of cell is follicular epithelial cell, and the second type of cell is parafollicular C cells. Differentiated Thyroid Cancer (DTC), which originates from follicular epithelial cells, is the primary form of thyroid gland cancer. Thyroid cancer is one of the endocrine malignant which has high recurrence rate after initial treatment. So, it is important to predict thyroid cancer recurrence efficiently. In this research, “Differentiated Thyroid Cancer Recurrence” dataset was collected from UCI machine learbning repository where 383 instances and 17 attributes were present. The target was to develop a machine learning (ML) model which can predict the Differentiated Thyroid Cancer Recurrence effectively and efficiently. This research developed a ML model by using Random Forest (RF) Classifier where it performed best among used other ML models. From evaluation metrics, Random Forest obtained the highest testing accuracy and also \(F_1\)-score where the testing accuracy was 98.70% and \(F_1\)-score was 97.30%. To evaluate further, this research used average cross-validation accuracy, precision, recall, and ROC-AUC curve, where average cross-validation accuracy on the training set was 96.08%.
Epilepsy Risk Prediction through Enhanced Machine Learning and Diverse Feature Selection Approaches
Abstract
Epilepsy is a prevalent neurological disorder affecting approximately 56 million people worldwide, yet early diagnosis remains challenging. Current diagnostic approaches are largely reactive, requiring patients to experience multiple seizures before confirmation, thereby limiting opportunities for preventive intervention. In this study, we developed a machine learning framework to predict epilepsy risk using clinical and demographic indicators. Using the Epilepsy Onset Prediction Dataset from Kaggle comprising 7,000 patient records with 20 features including demographic information, medical history, and clinical indicators, we evaluated nine machine learning algorithms. To address class imbalance inherent in medical datasets, we compared two sampling strategies: Synthetic Minority Over-sampling Technique (SMOTE) and random undersampling. All models underwent systematic hyperparameter optimization using GridSearchCV with 5-fold cross-validation. Our results demonstrate that oversampling substantially outperforms undersampling. Random Forest achieved an F1-score of 0.957 and ROC-AUC of 0.998 on the oversampled dataset. Furthermore, our stacking ensemble method achieved an F1-score of 0.984 and ROC-AUC of 0.999, representing approximately 6% improvement over existing state-of-the-art methods. The consistency between test and cross-validation metrics indicates robust generalization without overfitting. This study demonstrates that machine learning can effectively identify high-risk individuals for epilepsy, potentially enabling early intervention strategies that could benefit millions of patients worldwide.
Combating Misinformation: A Comparative Evaluation of Zero-Shot Models for Fake News Detection
Abstract
The number of fake content spread on the Internet platforms in rapid succession has led to the importance of the detection of fake news. One practical step is zero-shot classification with pre-trained language models that do not need further fine-tuning. In this paper, four transformer models DeBERTa, BART, ModernBERT, and CE-DeBERTa are evaluated on a smaller variant of the Kaggle Fake News dataset licensed under CC BY-4.0. Majority Voting and Soft Voting (with 0.85 and 0.92 thresholds) are also being used as ensemble learning methods to improve reliability. The highest accuracy (76.91%), the F1 score (72.64%), and the highest degree of uncertainty (59.40) were found in DeBERTa. However, ModernBERT showed less performance (F1 score: 62.40%) with the lowest uncertainty (31.09%). DeBERTa achieved the highest score of 0.6939 based on the composite performance score. The highest accuracy (70.23%) and the highest uncertainty (31.99%) were obtained with Soft Voting with a threshold of 0.92 among the ensemble methods. But Majority Voting turned out to have the strongest overall performance score (0.7089). In general, DeBERTa and the Majority Voting ensemble are rather successful in terms of trustworthy fake news identification.
Enhancing Financial Sentiment Analysis Using ML Models: A Comparative Study of Balanced vs. Imbalanced Datasets
Abstract
Financial sentiment analysis plays a crucial role in interpreting market sentiment from textual data, significantly influencing financial decision-making and forecasting. A common challenge in such analyses is the issue of class imbalance, where certain sentiments, such as negative or neutral sentiments, are underrepresented. The objective of this study is to assess the impact of dataset balancing on the performance of machine learning models in the context of financial sentiment classification. To achieve this, a comparative analysis was conducted using a diverse set of machine learning models, including Logistic Regression, Random Forest, AdaBoost, Gradient Boosting, XGBoost, SVM, KNN, Decision Tree, BERT, RoBERTa, and DeBERTa. These models were trained on both balanced and imbalanced datasets, employing resampling techniques to address class imbalance. The models were evaluated based on precision, recall, F1-score, and overall accuracy. The results revealed that balancing the dataset significantly improved the accuracy and reliability of the models, with the RoBERTa model achieving the highest accuracy on balanced datasets among all models considered. Interestingly, the SVM model demonstrated competitive performance, suggesting that non-transformer models can offer a time and resource efficient alternative while maintaining high accuracy. Furthermore, the findings suggest that exploring more advanced balancing techniques and deep learning methods may further improve the accuracy and robustness of financial sentiment classification.
DYL-Leaf: A Lightweight Distilled YOLO-based Model for Plant Leaf Disease Classification
Abstract
Agriculture is vital for global food security, but plant leaf diseases pose escalating threats, causing significant crop losses and economic damage. Traditional diagnostic methods are often time-consuming and resource-intensive, prompting the need for efficient, scalable solutions. This research addresses these challenges by proposing DYL-Leaf, a lightweight, distilled model designed for detecting 13 classes of potato, rice, and tomato leaf diseases from the PlantVillage dataset. Leveraging Knowledge Distillation (KD), a lightweight student model with only $\mathbf{5 4 5, 0 0 5}$ parameters, is trained to emulate a larger, custom YOLO-based teacher model (2.6 M parameters). The methodology optimizes both hard and soft losses using temperature-scaled KullbackLeibler divergence, ensuring the student model retains the teacher’s knowledge while being computationally efficient. Results demonstrate that the student model not only matches but surpasses the teacher’s performance, achieving a validation accuracy of 93.8% (vs. 92.9%), along with improved precision (94.00%), recall (93.23%), and F1-score (93.38%). Saliency maps were employed to interpret the model’s decision-making process, confirming its ability to focus on disease-specific features. These findings highlight the effectiveness of KD in creating lightweight, high-performing models for plant disease classification. By reducing computational requirements while maintaining accuracy, DYL-Leaf is well-suited for deployment in resourceconstrained agricultural environments.
Exploring the Barriers to Undergraduate Student Engagement with Research Papers: A Machine Learning and Statistical Analysis Approach
Abstract
University students often have significant obstacles to participating in research publications, motivated by characteristics such as language difficulties, approach problems, help from instructors, and content distribution options. This study discovered these characteristics and their influence on students' commitment to reading. A complete survey is done, concentrating on variables including text difficulty, accessibility problems, teacher encouragement, topic preferences, and study environment. To investigate the data, a multivariable linear regression model was applied to evaluate the connections between independent characteristics and student engagement. Preliminary studies have shown that the combination of these components has a significant impact on the participation decisions. Several advanced machine learning models such as Gradient Boosting Machines (GBM), Random Forest, and Support Vector Machines (SVM) were deployed. The Random Forest and GBM models achieved 95% accuracy in predicting student engagement. Machine Learning models were employed to identify and predict the key cognitive, motivational, and accessibility-related factors. This study emphasized the importance of solving these challenges to improve research skills, improve access, and establish a solid relationship between students and academic resources. In short, we can say that the results give practical advice to educational institutions to design a more effective approach to increase student interaction with research publications.
A Novel DEEP-SN Hybrid Approach for Segmenting and Classifying Solar Panel
Abstract
The detection of solar panels is vital for optimizing the efficiency of solar energy systems. This study presents a novel approach for solar panel detection using a hybrid deep learning model, DEEP-SN. The model combines the segmentation power of DeepLabV3 with the classification capabilities of the Swin Transformer. DeepLabV3 leverages a ResNet50 backbone and atrous spatial pyramid pooling (ASPP) to capture multi-scale information for precise image segmentation. With its ability to process long-range dependencies and spatial hierarchies, the Swin Transformer enhances the model’s classification accuracy. We employ a class-weighted loss function to handle a class imbalance in the dataset, and the OneCycleLR scheduler dynamically adjusts the learning rate during training. The model achieves a validation accuracy of 90%, a test accuracy of 90%, and an F1 score of 0.90, reflecting its strong performance in detecting solar panels. This approach offers a promising solution for automating solar panel detection, contributing to more effective monitoring and management of solar energy systems.
Predicting Emotional Well-Being from Social Media Usage: A Machine Learning Approach
Abstract
The rapid expansion of social media use has significantly influenced emotional well-being and has affected users in both positive and negative ways. The following paper examines the pattern of usage of social media, such as daily posts, likes, comments, time duration, etc., with respect to mental well-being through various machine learning models based on a data set collected from a verified Kaggle profile. Which was preprocessed by handling missing values, correcting misplaced features, and label encoding to train the models for effectively. Exploratory Data Analysis (EDA) has been carried out to find out the feature correlations that unveiled very important clues regarding user engagement with emotional well-being. Several Machine Learning algorithms, namely Decision Tree, Random Forest, LightGBM, CatBoost, and K-Nearest Neighbors are employed to predict emotional states based on user activity and usage patterns. Results confirm Random Forest as the best performing model, reaching test accuracy of $97.98 \%$ and validation accuracy of $98.67 \%$, thus proving quite robust in predicting emotional states from such data sets. These findings expose how sophisticated machine learning techniques can be proven effective toward an understanding of the complex relationship between social media behaviors and mental health and suggest further future potential applications in proactive support for mental health.
Predicting Academic Performance: Machine Learning Insights into GPA Determinants
Abstract
This study looks at predicting student GPA using machine learning models based on academic and non-academic data. Random Forest (RF), Gradient Boosting (GB), AdaBoost, Linear Regression (LR), Ridge, Lasso, Support Vector Regression (SVR), K-Nearest Neighbors (KNN), and XGBoost are among the models evaluated. R-squared ($\mathbf{R}^{\mathbf{2}}$), Mean Squared Error (MSE) and Mean Absolute Error (MAE) are used to assess performance. The most accurate models were Ridge and Linear Regression with accuracy of $95.68 \%$, which suggested a significant linear connection between variables and GPA. Support Vector Regression (SVR) and XGBoost also performed well but did not surpass the linear models. Because of its simplicity and accuracy, the research identifies Ridge Regression as the best model for predicting GPA, providing insightful information that may help educational institutions support student interventions.
Topic Modeling of Mpox-related Instagram Posts: Understanding Public Perception over Time
Abstract
Mpox is now a pandemic disease, and all its public health concerns have inspired very heated debates on Instagram. This paper is trying to capture possible public perceptions, concerns, and narratives through their Instagram posts on Mpox by applying two cutting-edge unsupervised topic models called BERTopic and Latent Dirichlet Allocation (LDA). A dataset of 60127 multilingual posts from July 2022 to October 2024 was processed and raised some clear thematic issues and vaccination hesitance, including disease awareness, awareness of misinformation, and geographic differences. BERTopic investigated nuanced location-specific issues such as humor on social media and technical risks; LDA gave a more extensive framework centered on broader topics like health emergencies and worldwide impact. The base of the analysis is a dataset of Mpox-related Instagram posts collected over a specified time. Pre-processing consists of text extraction and removal of URLs, tokenization, and removal of stop words, and also uses countVectorization. These findings serve to shed light on the interaction between public sentiment and health communication and the critical importance of customized outreach programs to counter misinformation and enhance awareness and public health response in outbreaks of infectious diseases. Findings yield insights into public health communication in setting proactive plans to curb misinformation and boost awareness, with future studies along the lines of cross-platform analysis as well as with multi-media data touching on the greater public perception.
Optimizing Alzheimer’s Disease Diagnosis Using Ensemble Machine Learning Techniques: A Comparative Study
Abstract
Alzheimer’s Disease (AD) remains one of the most challenging neurodegenerative disorders to diagnose due to the complexity of its underlying causes and progression patterns. The study involved Gradient Boosting, XGBoost, and Random Forests to achieve improvements in diagnostic accuracy. The study was performed on a dataset of 2,149 patient records featuring 35 attributes, showing that ensemble methods outperform conventional approaches in working with the complexity of medical data. The best model is Gradient Boosting showing the highest accuracy of 95%, which represents the possibility for the real-world application in medical diagnosis. The results enlighten about the role of preprocessing, hyperparameter tuning, and feature analysis to achieve reliability in prediction. Thereby, these findings contribute to the ever-accumulating compendium of machine learning underpinnings in medical research as they provide leads for future AD diagnostic frameworks.
Harvesting Insights: Unraveling Olive Dynamics and Climate Fluctuations Through Regression and SHapley Additive Explanations
Abstract
Exploring the historical evolution of olive cultivation, this research investigates the complex relationship between climate variations and olive dynamics. This work presents a comprehensive analysis of olive influx prediction based on a meticulously compiled dataset derived from climate change studies in the Levant. The 269 data points in the dataset include a variety of variables, including age prior to present, precipitation levels, and monthly temperatures. Through strategic pre-processing and model evaluation, the proposed ensemble approach emerges as the best model for predicting olive influx, outperforming other regression methods. Using SHAP analysis, key climate factors such as temperature, precipitation, and tree age are identified, shedding light on their importance to olive dynamics. Furthermore, this study goes beyond model performance, delving into a historical dataset to investigate the complex relationship between climate fluctuations and olive production. This research not only advances knowledge of the dynamics of olives in Tyre, Lebanon, but also offers a useful framework for other areas facing climate-related challenges in olive cultivation. This work bridges the gap between agricultural sustainability and machine learning, providing crucial insights for generating resilience in olive-growing regions addressing climate-related issues.
Optimizing Lithium-Ion Battery Recycling: Profits Unveiled Through Explainable AI and Transportation Mode Analysis
Abstract
Electric cars, renewable energy storage systems, and portable devices all depend on lithium-ion batteries. Concerns concerning efficient end-of-life management and recycling are brought up by the demand spike, nevertheless. This research investigates how to best recycle lithium-ion batteries by combining Explainable AI with transportation mode analysis. In this research, different machine learning models are explored for profit prediction in recycling operations including techniques like data acquisition, data processing, model selection, and Explainable AI frameworks. The most accurate model for predicting profits is the ensemble model, which combines the regressors from Random Forest and Extra Trees. It obtains the lowest MAE of 0.01, suggesting minimum variance between expected and actual earnings. The examination of transportation characteristics also sheds light on the best places for recycling and modes of transportation. Our findings reveal that combining ensemble model significantly improves profit prediction accuracy for lithium-ion battery recycling, with LIME identifying “Type” and “Cathode Scenario” as key predictive factors. The amalgamation of transportation mode analysis with Explainable AI offers a propitious methodology for optimizing the recycling of lithium-ion batteries, hence furnishing discernible insights into the decision-making process and augmenting operational efficiency.
Quantifying Climate Change Effects on Standard Minimum and Maximum Average Temperature Extremes in Bangladesh: A Machine Learning Regression Analysis from Past to Present
Abstract
The impact of climate change on temperature extremes is particularly significant in regions like Bangladesh, and it is a pressing global concern. To quantify the effects of climate change on Bangladesh’s standard minimum and maximum average temperature extremes, this study uses machine learning techniques on historical records from 1981 to 2010 with current data from 2022. By utilising information from the Statistical Yearbook Bangladesh 2022, which includes temperature readings from 44 different stations, this study offers a thorough evaluation of temperature fluctuations in various geographic areas. Several machine learning models, including Random Forest, Ridge Regression, Bayesian Ridge Regression, K-Nearest Neighbours Regression, and an Ensemble Model, are applied as part of the methodology. The following evaluation metrics are used: R-squared (\(R^{2}\)), Mean Absolute Percentage Error (MAPE), Root Mean Squared Error (RMSE), Mean Squared Error (MSE), and Root Mean Squared Logarithmic Error (RMSLE). The study’s conclusions highlight notable yearly temperature variations and dynamic patterns of climate change in important places like Khulna, Chattogram, Teknaf, and Dhaka. The outcomes show that, compared to other models, the Ensemble Model performs better across evaluation metrics and is the most successful at predicting temperature extremes. Furthermore, comparing the actual and predicted temperatures for a few selected stations reveals significant deviations that underscore Bangladesh’s changing climate dynamics. This research enhances our understanding of the impact of climate change on temperature extremes in Bangladesh. This study helps stakeholders, researchers, and policymakers better understand the difficulties brought on by the region’s changing climate patterns.
Optimized Machine Learning Models and Stacking Hybrid Approach for Chronic Kidney Disease Prediction
Abstract
Chronic Kidney Disease (CKD) is a severe disease that requires being diagnosed at an early stage to stop potential complications. Machine learning methods show an effective solution for predicting CKD, which enables healthcare professionals to make timely decisions. However, the selection of proper models along with their performance optimization remains the key essential challenge. This study examines the effectiveness of several machine learning algorithms, such as k-nearest neighbours (KNN), decision tree classifier (DT), logistic regression, naïve bayes (NB), and gradient boosting (GB), in predicting CKD cases. In addition, a stacking hybrid model focuses on enhancing classification effectiveness as well as model robustness for preventing overfitting. Each baseline model undergoes hyperparameter tuning for maximizing accuracy through optimal setting determination. Performance evaluation of the models uses multiple measurement criteria that include accuracy, precision, recall, f1-score, Jaccard score, and AUC-ROC curve, where the optimized model and hybrid model have shown better results. These metrics provide a complete insight into how well models classify CKD and non-CKD cases correctly. The results demonstrate that each optimized model improves model performance by using hyperparameter tuning. Moreover, the stacking hybrid model (using 3 of the best optimized models) performs better than each of the optimized classifiers, showing that the potential of integrating various models can boost the ability of the prediction. This study highlights the optimized model and a well-tuned hybrid model can be a valuable approach for enhancing diagnostic accuracy in medical decision-making, aiding in the early identification and management of CKD. However, The source code of this study is available at this link: https://github.com/dragontea72/Chronic-Kidney-Disease-Prediction-.
Explainable AI in Feature Selection: Improving Classification Performance on Imbalanced Datasets
Abstract
Particularly in biomedical applications, feature selection plays a critical role in enhancing the interpretability and efficacy of machine learning models. This work examines the performance of Explainable AI (XAI), Information Gain (IG), and Principal Component Analysis (PCA) methods on an imbalanced dataset pertaining to stroke prediction. Data from 4,603 patient records, including 362 instances of stroke from the National Health and Nutrition Examination Survey are used in this investigation. Methodologically, IG is used for feature ranking, PCA is used to reduce dimensionality and XAI techniques are used to improve model transparency. The chosen features are used to assess the performance of several machine learning models, including Random Forest, Support Vector Machine, k-Nearest Neighbours, and Logistic Regression, in terms of classification. Our experimental results show that the combined PCA-IG approach significantly enhances classification accuracy, achieving 91.75%. Furthermore, LIME-based feature selection outperformed in precision, recall, and F1 score, with the highest accuracy at 91.86%. LIME discovered nine positive impact features, highlighting the top contributors in the dataset. We also applied the same feature selection technique to datasets from other domains. These findings highlight the robustness of using PCA-IG and XAI approaches separately to create reliable and understandable machine learning models for healthcare and other applications. By offering insights into the optimal use of PCA, IG, and XAI to enhance the accuracy and practicality of machine learning models in healthcare and other domains, this paper advances the field of feature selection across all areas of data analysis.
Explainable Federated Stacking Models with Encrypted Gradients for Secure Kidney Medical Imaging Diagnosis
Abstract
In the field of medical data analysis, both privacy retention and model interpretability are of utmost importance. This paper focuses on the vulnerability of kidney medical image analysis through Federated Learning (FL) with privacy-preserving measures and explainable artificial intelligence (XAI). We introduce a new set of proposals termed Federated Stacking Fusion with Encryption (FedStackEncFL), which integrates the predictive capabilities of ResNet-101 and InceptionV3 architectures through a stacking fusion technique. This approach also guarantees the quality and reliability of the features extracted while preserving the federated data’s privacy by encrypting it through Cheon-Kim-Kim-Song (CKKS) during Federated Averaging (FedAvg). Moreover, to optimize privacy, during the computation, the method of adding Gaussian gradient noise is used. The efficiency of the developed technique is proven by the complex experimental data proving the efficiency of the humanitarian approach to enhance indicators of model precision, recall, F1-score, and accuracy in Kidney disease diagnosis. Furthermore, the integration of XAI techniques like GradCAM, GradCAM++, and ScoreCAM provides insightful visual explanations, enhancing the interpretability of our model’s predictions in medical diagnostics.
From Theory to Reality: Tracing the Milestones of Quantum Information Systems
Abstract
This paper provides a comprehensive historical exploration of the development of quantum information systems (QIS), tracing the foundational concepts and seminal discoveries that have significantly advanced this revolutionary field.Starting with early 20thcentury breakthroughs in quantum mechanics, it highlights key milestones such as the elucidation of the EPR paradox, the introduction of quantum error correction, and the development of groundbreaking algorithms like Shor's and Grover's.The paper also addresses notable experimental achievements that have transitioned QIS from theoretical frameworks to practical applications, including quantum teleportation and fault-tolerant quantum computation.Additionally, the manuscript discusses current limitations and research challenges, including the fragility of qubits, error correction complexities, scalability issues, and the need for advanced quantum algorithms.The potential applications of QIS in various fields such as pharmaceuticals, financial services, climate modeling, energy, healthcare, defense, and manufacturing are explored, demonstrating its far-reaching implications.By chronicling the remarkable progress and future opportunities in QIS, this paper offers valuable insights into the transformative potential of quantum technologies for information processing, cryptography, and beyond.
Uncovering Hidden Realities of Child Labor Abuse in Egyptian Workplaces with Machine Learning and Explainable AI
Abstract
Child labor abuse is a pervasive global problem, especially in Egyptian workplaces, where its bureaucratic processes make it difficult to detect and intervene such phenomena.This research presents an in-depth investigation combining explainable artificial intelligence (XAI) and machine learning (ML) to uncover the hidden reality of child labor exploitation in Egyptian companies.Applying six different machine learning models and analyzing, on a published dataset of mistreatment of employed children in Egypt, Artificial Neural Network (ANN) model turns out to performs better than the other models.It shows improvements in accuracy, precision, recall, F1-score, and AUC, over other models including Random Forest, Support Vector Classifier (SVC), and AdaBoost, with an accuracy of 96.00%.Additionally, Explainable AI methods like SHAPASH analysis offer insightful information about the variables, such as working sector and night shift work, driving abuse predictions.With its data-driven method to reveal hidden realities, this research not only promotes a safer environment for kids in Egyptian workplaces, but also these study findings can be expanded in identifying other form of abuses such as elder abuse, substance abuse and others.
A Comprehensive Approach to Gold Price Prediction Using Machine Learning and Time Series Models
Abstract
This paper provides a comprehensive analysis of gold rate fluctuations in the USA and introduces an innovative method for more accurately predicting gold price changes using advanced machine-learning techniques. The dataset consists of daily gold rate data from January 1, 1985, to September 8, 2023, expressed in this country’s national currency. However, for our research, we specifically utilized the data expressed in United States Dollars (USD). Several machine learning models, ARIMA, LSTM, FB Prophet, SVM, Random Forest, and XGBoost, were employed to model and predict these time series data. The predictive performance of these methods is compared, providing insights into their effectiveness in capturing non-linear and long-term trends in gold prices. The outcomes reveal how machine learning could guide financial time series forecasting better and bring to light the most powerful clues about the way gold prices are moving on the global stage.
Machine Learning Models for Predicting Olympic Medal Outcomes
Abstract
In this paper, the prediction of Olympic medal winners has been explored using various machine learning models, utilizing a dataset spanning 128 years of Olympic history from 1896 to 2024. Thirteen Machine Learning models - Logistic Regression, Polynomial Logistic Regression, XGBoost, Random Forest, KNN, Naive Bayes, LightGBM, AdaBoost, Decision Tree, Extra Trees, Gradient Boosting, Neural Network, and SVM were used and evaluated. The dataset was preprocessed, and models were trained and tested to predict future medal outcomes. Superior performance metrics were demonstrated by ensemble models such as XGBoost, LightGBM, and Gradient Boosting with accuracy rates of 83%, 84%, and 84% respectively, and notable AUC values. Conversely, lower performances were exhibited by simpler models like Logistic Regression, Polynomial Logistic Regression, and SVM. Discrepancies in prediction graphs suggested potential issues in model training or dataset encoding. Future research will focus on integrating additional features and more sophisticated techniques to enhance model performance and prediction accuracy. The findings are anticipated to contribute significantly to the field of sports analytics and assist national Olympic committees in strategic planning and resource allocation for future Olympic Games.
From Data to Diagnosis: Applying Machine Learning Model for Reliable Heart Disease Prediction
Abstract
Heart disease remains a significant global health challenge, demanding precise and interpretable predictive tools to aid early intervention and improve patient outcomes. This study applies the XGBoost machine learning algorithm to develop a robust model for predicting cardiovascular disease risk. Leveraging a comprehensive dataset that includes demographic, lifestyle, and clinical features, our approach combines extensive data preprocessing, feature engineering, and Bayesian hyperparameter tuning to achieve superior predictive accuracy, with an AUC of 0.98 and an accuracy of 97.48%. Key predictive factors, identified through SHAP (SHapley Additive exPlanations) values, include cholesterol, blood pressure, and physical activity, offering actionable insights for clinicians. The model’s interpretability supports its potential as a reliable, real-time diagnostic tool, enhancing clinical decision-making and risk assessment. This research advances predictive healthcare analytics by showcasing the effectiveness of sophisticated machine learning methods, including XGBoost, for precise and interpretable heart disease forecasting. Subsequent study will seek to validate the model across varied populations and incorporate it into clinical environments for adaptive, individualized cardiovascular risk management.
Advancing Mental Health Problems with Machine Learning and Genetic Algorithms for Anxiety Classification in Bangladeshi University Students
Abstract
Mental health challenges, particularly anxiety, are a growing concern among university students in Bangladesh, impacting both their well-being and academic performance. This study aims to address this growing issue by developing a robust machine learning model tailored to classify anxiety levels among students. We employed eight well-known machine learning algorithms, with particular emphasis on a customized Genetic Algorithm (GA) optimized Logistic Regression (LR) model. The models were rigorously trained and evaluated using 5-fold cross-validation on a newly obtained mental health dataset that had never been investigated with such advanced approaches. This dataset contains information on 1,977 students from 15 of Bangladesh’s leading universities and was thoroughly examined by five academics with 10 to 20 years of experience in academia and research. Our findings indicate that classic models, such as the Random Forest Classifier and Support Vector Classifier, attained accuracies of 92.68% and 96.72%, respectively. However, our proposed GA-based LR model succeeded them all, not only in accuracy but also in precision, recall, and F1-score, with an impressive 99.49% accuracy. The study also identified key anxiety-inducing factors, such as academic pressure and worry about academic affairs, providing valuable insights for targeted mental health interventions. These findings indicate the effectiveness of our customized GA-based ML model in enhancing mental health evaluations by identifying the fundamental causes of anxiety, providing vital insights to help Bangladeshi university students maintain their mental health.
Analysing the Efficacy of UML in Explaining Object-Oriented Concepts to Undergraduate Computer Science Students
Abstract
This study used the undergraduate course Object-Oriented Analysis and Design as a case study to assess how the depth of knowledge of object-oriented concepts was developed in preparation for learning the Unified Modeling Language (UML). Two surveys were conducted to assess student learning. The first survey identified the concepts of object orientation that students should be aware of while taking object-oriented courses. The second survey examined the depth of knowledge gained from learning UML. The results show that teaching UML significantly increases students' knowledge of UML compared to just learning it in object-oriented courses. This study concluded with reflections on the effectiveness of UML as an approach to developing the concept when taught at the undergraduate level.
Evaluating the Effectiveness of ML Algorithms in Water Potability Prediction Aligned with SDGs
Abstract
This research addresses the critical need for accurate water quality prediction, aligning with Clean Water and Sanitation (SDG 6) and Good Health and Well-being (SDG 3). Utilizing a diverse dataset of water quality attributes, we developed and evaluated multiple machine learning models to predict water potability, contributing to SDG 9 (Industry, Innovation, and Infrastructure) through technological advancement. Our primary contribution is a Hyperparameter-Tuned Histogram Gradient Boosting Classifier, which outperformed several other algorithms, achieving an accuracy of 67%, a ROC-AUC score of 70%, and an F1-score of 66%. The study highlights challenges of class imbalance and feature distribution in water quality prediction, addressing aspects of SDG 11 (Sustainable Cities and Communities) by improving urban water management. Our findings have significant implications for water resource management, supporting SDG 12 (Responsible Consumption and Production) through efficient resource utilization. While acknowledging limitations such as single dataset reliance, this research combines modern machine learning with comprehensive analysis to enhance water quality assessment. It contributes to SDG 17 (Partnerships for the Goals) by providing a framework for collaboration between technology and environmental sectors. Future directions emphasize improving model performance and real-world applicability, potentially supporting SDG 13 (Climate Action) by enhancing resilience to water-related climate impacts.
Explainable Detection and Analysis of Cauliflower Leaf Diseases
Abstract
This study presents a novel approach for detecting diseases in cauliflower using advanced deep-learning techniques. Based on the scope of neural networks such as NasNet Mobile and InceptionV3, cauliflower’s distinct leaf diseases were first detected and then classified in this study. To generate the dataset, pictures were taken under various weather conditions. The integration of NasNet Mobile and InceptionV3 not only improved the detection rate but also increased computational efficiency. This strategy will immensely help if employed during practical activities such as farming. The results elicited the use of a deep learning model to improve existing plant disease diagnosis techniques. It applied of LIME (Local Interpretable Model-agnostic Explanations) and Grad-CAM (Gradient-weighted Class Activation Mapping) as explainable artificial intelligence techniques to enhance the understanding of the suggested model’s output and, therefore, foster decision-making transparency. In terms of the suggested model’s performance metrics, we got a precision of 99.80%, a recall of 100% and F1-score of 99.90%. This innovative method, with its high accuracy and speed, holds promise for widespread adoption by farmers to maintain plant health and optimize crop yields.
Multicriteria Decision Analysis for Optimal Internet Service Provider Selection Using Calibrated Random Forest
Abstract
The Internet is integral to modern life, with Internet service provider (ISP) offering appealing deals to meet the demand for unlimited data. However, reality often falls short of expectations. While recommendation systems exist, user-centric options are rare. This paper proposes a novel ISP selection methodology using user experience data and a calibrated random forest (CRF) model. Unlike traditional methods that focus on advertised features, this approach emphasizes user-defined criteria such as cost, device connectivity, and technical support experience. By analyzing survey data, the model highlights the critical link between user needs and support quality, enabling users to choose ISPs that prioritize customer service. The model demonstrates promising results with a strong R-squared value and low mean squared error (MSE). This user-centric approach fosters informed decision-making, potentially driving competition and encouraging ISPs to improve service standards, laying a foundation for future developments in ISP selection.
Predictive Analysis of Internship and Job Placement Success in Computer Science Education
Abstract
The growing competition in the job market calls for a more in-depth comprehension of the factors that affect successful internship and job placement results in computer science education. This study provides a thorough examination of 25 machine learning classification models in order to forecast the job prospects of computer science students. Out of these models, the Random Forest Classifier stood out as the best performer, with an accuracy, precision, recall, and F1 score of 82.90%, 82.90%, 100.00%, and 90.65% respectively. The research investigates the influence of different factors like CGPA, course completion, specialization, project attachments, and extracurricular activities on job placement outcomes. Moreover, the study highlights the significance of model calibration, leading to enhanced performance of the Random Forest Classifier. The results provide useful information for schools, enabling them to provide focused assistance for students and improve their curriculum. Future efforts will be directed towards enhancing the model, incorporating current industry data in real-time, and experimenting with advanced methods to enhance the accuracy of predictions. This research seeks to enhance predictive analytics in the education sector, specifically by enhancing placement results for computer science graduates.
Predictive Modeling and Early Detection of White Spot Disease in Shrimp Farming Using Machine Learning: A Case Study in Bangladesh
Abstract
This paper addresses the severe threat of White Spot Disease (WSD) to the shrimp farming industry in Bangladesh and proposes a comprehensive methodology for predictive modeling and early detection using advanced machine learning techniques. The study emphasizes the importance of shrimp aquaculture in Bangladesh, outlines the impact of WSD outbreaks, and sets objectives for developing a specific prediction model for shrimp farming. Employing Random Forest, Multinomial Naive Bayes, and Bagging with Decision Trees, the research achieves impressive accuracy rates exceeding 90%, with the Bagging with Decision Trees classifier and Random Forest reaching a maximum accuracy of 97.87%. The outcomes hold significant implications for early WSD identification in the shrimp farming industry, contributing to Bangladesh’s GDP and providing valuable insights for managing aquatic. The disease outbreaks globally.
Evaluating Teachers' Performance through Aspect-Based Sentiment Analysis
Abstract
This research demonstrates a novel approach for evaluating teacher performance by conducting aspect-based sentiment analysis (ABSA) on student feedback. A large dataset of over 2 million student comments about teachers is analyzed using cutting-edge natural language processing and customized deep learning techniques. The methodology involves identifying positive, negative and neutral aspects of teaching using a BiLSTM model. Rigorous preprocessing, domain adaptation, and performance metrics ensure a robust and objective evaluation. The granular, nuanced insights obtained through this aspect-level sentiment analysis enable educational institutions to provide targeted and unbiased feedback to teachers on their strengths and areas needing improvement. Moreover, this work lays the foundation for detecting potentially fraudulent reviews in academic settings – a crucial capability for safeguarding assessment integrity. The detailed aspect-based analysis methodology presented here significantly advances subjective and holistic evaluation practices. This research has far-reaching implications for enriching teacher development while upholding the credibility of performance assessments through sentiment analysis innovations.
Sustainability-Driven Hourly Energy Demand Forecasting in Bangladesh Using Bi-LSTMs
Abstract
This research presents a comprehensive study on developing and evaluating a deep learning-based forecasting model for hourly energy demand prediction in Bangladesh. Leveraging a novel dataset obtained from the Power Grid Company of Bangladesh (PGCB), the proposed model utilizes bi-directional long short-term memory networks (Bi-LSTMs), implemented through Tensor-Flow and Keras libraries. The study meticulously preprocesses the data, handling missing values and ensuring compatibility with the selected models. The models are trained and evaluated using Mean Absolute Error (MAE) and Mean Squared Error (MSE) metrics, revealing promising results of 376.72 of MAE. The experimental findings demonstrate the effectiveness of the developed forecasting model, showcasing its capability to predict energy demand accurately. The insights derived from this study pave the way for enhanced energy management strategies, fostering sustainable and efficient energy utilization practices.
Understanding the Dynamics of Dengue in Bangladesh: EDA, Climate Correlation, and Predictive Modeling
Abstract
Dengue, a mosquito-borne viral infection, poses a significant threat, especially in warm, tropical climate countries like Bangladesh, India, Thailand, Malaysia, Laos, etc. This study is solely focused on the dengue data of Bangladesh as it explores the historical dengue data spanning 23 years (2000 to 2022) for EDA purposes, with a focus on 9 years (2014–2022) divisional data for model performance analysis. Additionally, climate data was collected for the same period to examine the potential correlation between dengue cases and climate factors. Machine learning (ML) and Deep learning (DL) models, including Random Forest Regression (RFR), Long Short-Term Memory (LSTM), and LSTM with Artificial Neural Networks (ANN), were implemented and validated against ground truth data. The results reveal notable differences in performance between ML and DL models when handling imbalanced datasets with outliers, with RFR outperforming LSTM when compared to the ground truth data. The study uncovers significant correlations between dengue cases and climate factors like humidity, temperature, and precipitation. The insights gained from this research have practical implications for dengue prevention and control efforts in Bangladesh and beyond, paving the way for more effective strategies and interventions.
Target and Precursor Named Entities Recognition from Scientific Texts of High-Temperature Steel Using Deep Neural Network
Abstract
Named Entity Recognition (NER) is an essential task in natural language processing, especially in the domain of scientific texts. This paper presents a study of NER for scientific texts in high-temperature steel, a type of alloy used in various applications where high temperatures prevail. We propose a NER system using Bi-LSTM with a domain-specific embedding approach and evaluate its performance on a test dataset. The study results show that the proposed NER system achieves an F1 score of 0.99, indicating that it can accurately identify and classify named entities in scientific texts about high-temperature steel with high precision and recall. The proposed approach was more effective than the classical machine learning-based approach. Our results suggest that the domain-specific embedded Bi-LSTM technique can be an effective approach for NER in scientific texts, especially in specialized domains such as high-temperature steel.
Medical Named Entity Recognition (MedNER): A Deep Learning Model for Recognizing Medical Entities (Drug, Disease) from Scientific Texts
Abstract
Medical Named Entity Recognition (MedNER) is an indispensable task in biomedical text mining. NER aims to recognize and categorize named entities in scientific literature, such as genes, proteins, diseases, and medications. This work is difficult due to the complexity of scientific language and the abundance of available material in the biomedical sector. Using domain-specific embedding and Bi-LSTM, we propose a novel NER model that employs deep learning approaches to improve the performance of NER on scientific publications. Our model gets 98% F1-score on a curated data-set of Covid-related scientific publications published in multiple web of science and pubmed indexed journals, significantly outperforming previous approaches deployed on the same data-set. Our findings illustrate the efficacy of our approach in reliably recognizing and classifying named entities (drug and disease) in scientific literature, opening the way for future developments in biomedical text mining.
Impact of COVID-19 Lockdowns on Air Quality in Bangladesh: Analysis and AQI Forecasting with Support Vector Regression
Abstract
Over the past few decades, air pollution has emerged as a significant environmental hazard, causing premature deaths in Southeast Asia. The proliferation of industrialization and deforestation has resulted in an alarming increase in pollution levels. However, the COVID-19 pandemic has significantly reduced the amount of volatile organic compounds and toxic gases in the air due to the decrease in human activity caused by lockdowns and restrictions. This study aims to investigate the air quality in various geographical areas of Bangladesh, comparing the air quality index (AQI) during different lockdown periods to equivalent eight-year time spans in 10 of the country’s busiest cities. This study demonstrates a strong correlation between the rapid and widespread dispersion of COVID-19 and air pollution reduction in Bangladesh. In addition, we evaluated the performance of Support Vector Regression (SVR) in AQI forecasting using the time series dataset. The results can help improve machine learning and deep learning models for accurate AQI forecasting. This study contributes to developing effective policies and strategies for reducing air pollution in Bangladesh and other countries facing similar challenges.
Material Named Entity Recognition (MNER) for Knowledge-Driven Materials Using Deep Learning Approach
Abstract
The scientific literature contains an abundance of cutting-edge knowledge in the field of materials science, as well as useful data (e.g., numerical values from experimental results, properties, and structure of materials). To speed up the identification of new materials, these data are essential for data-driven machine learning (ML) and deep learning (DL) techniques. Due to the large and growing amount of publications, it is difficult for humans to manually retrieve and retain this knowledge. In this context, we investigate a deep neural network model based on Bi-LSTM to retrieve knowledge from published scientific articles. The proposed deep neural network-based model achieves an F1 score of \(\widetilde{9}7\)% for the Material Named Entity Recognition (MNER) task. The study addresses motivation, relevant work, methodology, hyperparameters, and overall performance evaluation. The analysis provides insight into the results of the experiment and points to future directions for current research.
Predicting Carboxymethyl Cellulase assay (CMCase) production using Artificial Neural Network and explicit feature selection approach
Abstract
This paper presents a method for predicting carboxymethyl cellulase (CMCase) production using artificial neural networks (ANNs) and an explicit feature selection approach. A dataset of CMCase production experiments was collected, and an explicit feature selection approach was applied to select the most relevant features for CMCase production prediction. The ANN model was trained using both the selected features and all available features of the CMCase production data. The results showed that the explicit feature selection approach improved the performance of the ANN model in terms of prediction accuracy compared to using all the features available in the dataset. The main effect analysis (MEA) was found to be the best method for selecting the explicit features for predicting CMCase production. The ANN model trained using the MEA identified features, achieved 96.3% R2score and a MAE of 0.057 and a MSE of 0.035. The proposed method is an effective approach for predicting CMCase production and can be used to optimize CMCase production and reduce costs in various industries.
Predicting the Success of Suicide Terrorist Attacks using different Machine Learning Algorithms
Abstract
Extremism has become one of the major threats throughout the world over the past few decades. In the last two decades, there has been a sharp increase in extremism and terrorist attacks. Nowadays, terrorism concerns all nations in terms of national security and is considered one of the most priority research topics. In order to support the national defense system, academics and researchers are analyzing various datasets to determine the reasons behind these attacks, their patterns, and how to predict their success. The main objective of our paper is to predict different types of attacks, such as successful suicide attacks, successful non-suicide attacks, unsuccessful suicide attacks, and unsuccessful non-suicide attacks. For this purpose, various machine learning algorithms, namely Random Forest, K Nearest Neighbor, Decision Tree, LightGBM Boosting, and a feedforward Artificial Neural Network called Multilayer Perceptron (MLP), are used to determine the success of suicide terrorist attacks. With an accuracy rate of 98.4% and an AUC-ROC score of 99.9%, the Random Forest classifier was the most accurate among all other algorithms. This model is more trustworthy than previous work and provides a useful comparison between machine learning methods and an artificial neural network because it is less dependent and has a multiclass target feature.
An automated monitoring and environmental control system for laboratory-scale cultivation of oyster mushrooms using the Internet of Agricultural Thing (IoAT)
Abstract
This research paper presents an automated system for controlling and monitoring the cultivation of oyster mushrooms in a laboratory facility. The system uses an Internet of Things (IoT) based approach to automate the entire process. The main objective of this system is to make indoor mushroom cultivation easier and cost effective. The proposed system has proved to be very effective in saving labor cost by automatically monitoring and controlling the environment. Moreover, the proposed system has a real time update and monitoring feature which helps the grower to take immediate action. This paper provides a detailed insight into the project, including the development process, automation system, hardware setup, technical specifications, and results achieved after successful implementation.
A Trend Analysis of crimes in Bangladesh
Abstract
This paper presents a trend analysis of crimes in Bangladesh based on the data provided by Bangladesh Police for the last ten years, from 2010 to 2019. The data contains the number of different crimes organized by the police units in different areas of Bangladesh. The aim of this work is to analyze the crimes committed in the last ten years, establish relationships between the different types of crimes, and identify patterns between them. In order to analyze the data, the dataset provided by Bangladesh Police is transformed into different forms, and then the correlation between different features is determined. The four most important features or crime types, namely murder, narcotics, smuggling, and dacoity, are analyzed in depth by analyzing the significance of the features. This study presents the overall scenario of the crime types and rates observed in the last decade and could be of great help to the law enforcement agencies to make future decisions based on the analysis presented in this paper.
An adaptive Medical Cyber-Physical System for post diagnosis patient care using cloud computing and machine learning approach
Abstract
Medical care is one of the most basic human needs. Due to the global shortage of doctors, nurses, and other healthcare personnel, medical cyber-physical systems are quickly becoming a viable option. Post-diagnosis surveillance is an essential application of these systems, which can be performed more successfully using various monitoring devices rather than active observation by nurses in their physical presence. However, most existing solutions for this application are rigid and do not consider current difficulties. Intelligent and adaptive systems can overcome the challenges because of the advances in relevant technology, especially healthcare 4.0. Therefore, this work presents an adaptive system based on cloud and edge computing architecture and machine learning approaches to perform post-diagnosis medical tasks on patients, thus reducing the need for nurses, especially in the post-diagnosis phase.
A Behavioral Trust Model for Internet of Healthcare Things Using an Improved FP-Growth Algorithm and Naïve Bayes Classifier
Abstract
Healthcare 4.0 has revolutionized the delivery of healthcare services during the last years. Facilitated by it, many hospitals have migrated to the paradigm of being smart. Smartization of hospitals has reduced healthcare costs while providing improved and reliable healthcare services. Thanks to the Internet of Healthcare Things (IoHT) based healthcare delivery frameworks, integration of many heterogeneous devices with varying computational capabilities has been possible. However, this introduced a number of security concerns as many secure communication protocols for traditional networks can not be verbatim employed on these frameworks. To ensure security, the threats can largely be tackled by employing a Trust Management Model (TMM) which will critically evaluate the behavior or activity pattern of the nodes and block the untrusted ones. Towards securing these frameworks through an intelligent TMM, this work proposes a machine learning based Behavioral Trust Model (BTM), where an improved Frequent Pattern Growth${\left(iFP-Growth\right)}$algorithm is proposed and applied to extract behavioral signatures of various trust classes. Later, these behavioral signatures are utilized in classifying incoming communication requests to either trustworthy and untrustworthy (trust) class using the Naïve Bayes classifier. The proposed model is tested on a benchmark dataset along with other similar existing models, where the proposed BMT outperforms the existing TMMs.
Comparison of document similarity algorithms in extracting document keywords from an academic paper
Abstract
The idea of this study is to validate a list of keywords derived from a scientific article by a domain expert from years of knowledge with prominent document similarity algorithms. For this study, a list of handcrafted keywords generated by Electric Double Layer Capacitor (EDLC) experts are chosen, and relevant documents to EDLC are considered for the comparison. Then, different similarity calculation algorithms were employed in different settings on the documents such as using the whole texts of the documents, selecting the positive sentences of the documents, and generating similarity score with automatically extracted keywords from the documents. The experiment’s outcome provides us with findings that the machine-generated keywords are mostly similar to the curated list by the domain experts. This study also suggests the preferable algorithms for similarity calculation and automated key-phrase extraction for the EDLC domain.
Awareness to Deepfake: A resistance mechanism to Deepfake
Abstract
The goal of this study is to find whether exposure to Deepfake videos makes people better at detecting Deepfake videos and whether it is a better strategy against fighting Deepfake. For this study a group of people from Bangladesh has volunteered. This group were exposed to a number of Deepfake videos and asked subsequent questions to verify improvement on their level of awareness and detection in context of Deepfake videos. This study has been performed in two phases, where second phase was performed to validate any generalization. The fake videos are tailored for the specific audience and where suited, are created from scratch. Finally, the results are analyzed, and the study’s goals are inferred from the obtained data.
Performance comparison of HTTP/2 for Common E-Commerce Web Frameworks with Traditional HTTP
Abstract
Abstract HTTP/2 is a cutting-edge Web convention predicated on Google’s SPDY convention which tries to tackle the deficiencies and rigidity of HTTP/1. As e-commerce websites have become a significant medium for Online shopping, this paper demonstrates that whether HTTP/2 can authentically avail the performance of an e-commerce web browsing over HTTP or not. This paper states that we have studied about the HTTP/2 Implementation & performance analysis for prevalent web frameworks, where we have culled two different e-commerce web frameworks Laravel, WordPress (WooCommerce). At first, we have implemented two e-commerce sites in HTTP, then we have implemented those into HTTP/2. We additionally deployed them on the live server. By utilizing the Webserver Stress Tool & Selenium Web Driver, we have evaluated the performance under the sundry network environments. Selenium has given better results among all. But the webserver stress tool has exhibited some errors only for the http/2. This is our only constraint for this work. But still, we are endeavoring to resolve the problem. Simulation results have shown how HTTP/2 has influenced the page load time in our e-commerce websites. After analyzing all the simulation results, we have decided that Laravel is a more superior e-commerce web framework for both protocols, especially for HTTP/2.
A Geofencing-based Recent Trends Identification from Twitter Data
Abstract
Abstract For facilitating users from information overloading by finding recent trends in twitter, several techniques are proposed. However, most of these techniques need to process extensive data. Therefore, in this paper, a geofencing-based recent trends identification technique is proposed, which acquires data based on a geofence. Afterwards, they are cleaned and the weight of these tweet data is calculated. For that, the frequency of tweet texts and hashtags are taken into account along with a boosting factor. Thereafter, they are ranked to recommend recent trends to the user. This proposed technique is applied in developing a system using Java and python. It is compared with other relevant systems, where it demonstrates that the performance of the proposed system is comparable. Over and above, since the proposed system integrates geofencing feature, it is more preferable over other systems.
Location, Context and Device aware Framework (LCDF): A unified Framework for Mobile Data Management
Abstract
The objective is to propose an unified framework to manage and optimize the data being generated by the mobile devices we regularly use, specially the smart cell phones, tablet pc, wearable devices after studying the current practiced frameworks. This study proposes a top-level framework for data consuming and sharing between mobile devices as well as the data collection and data management methods. This study has addressed some issues with the current data management frameworks and proposed the way how those issues can be avoided in future.
IoT Based Healthcare Middleware
Abstract
In recent times, personal medical information can be of great value in times of crisis. Persistent medical information and history can be lost during transitions or carelessness. The aim of this work is to propose a middleware that shall help transfer medical information of patient to doctors and health care facilities. It is designed to function and work as a bridge between different healthcare services. The functions used by the middleware are driven by means of Internet of things (IoT). On the situation used by us cellular phones and personal computers will be used to exchange and alter information. Each patient's individual social security number (SSN) or any other equivalent identification will be their patient-ID making it a more simple way to organize each patient's unique database which shall be enriched with their medical history from the time they are born. Data will be stored in cloud storage as well as embedded in the patient's device to ensure swift data retrieval. Data will be transferred from one device to another through web applications and mobile applications. The system will significantly decrease the communication gap between patients and doctors while guaranteeing superior data exchange between all parties involved. During emergency, it will first send basic information regarding the patient to the nearest ambulances. Secondly it will alert nearest hospitals, and thirdly, allowing users to send detailed and personal data to a specific doctors to ensure precise and accurate treatment.
The issues and the possible solutions for implementing self-driving cars in bangladesh
Abstract
The main purpose of this study is to find the issues and suggest possible solutions regarding the implementation of Self-driving cars on Bangladeshi roads. A Self-driving car is fairly in its infancy today, but once implemented this can revolutionize our traffic management and transport system. This paper addresses some of the hurdles the technology might face in the country and offers a few measures that can be taken to overcome the downsides.
SHIMPG: Simple Human Interaction with Machine using Physical Gesture
Abstract
With all the advances accomplished in the computer gaming world progressively emerging more equipment's into the market in order to communicate with computers. Most of these sophisticated devices are for targeted applications or special purpose and also very expensive. With this in our mind we propose a Simple Human Computer Interaction with Machine using Physical Gesture Framework which uses regularly used equipment like low resolution camera and also has a very robust uses like simple games and smart house. The main goal of this paper is to outline a mechanism of computer vision for controlling any application or hardware. For this purpose we propose a generalized framework which can be seamlessly integrated with any networked controlled application as it works as a separate engine which can be easily integrated. In addition we have prototyped a simple application using this engine, image processing and 3D model augmentation.
Cricket Shot Classification Using Motion Vector
Abstract
Cricket shots cannot be detected yet from single video sample without multiple view camera and other tools like sonar, speedometer. Extracting salient feature and optical flow from videos of cricket shots is still a challenge. In cricket, body parts movement created several different directional optical flows. So we propose motion estimation approach related to classifying the shots using 3D MACH for action recognition. Our methodology defines 8 classes of angle ranges to detect cricket shots. Our method is grounded on Motion vectors that help to measure the angle of any precise cricket shot. An adequate accuracy level for the shots is established for this particular approach.
Optimized Tracking using VAIT for Moving Objects
Abstract
This Automated object tracking and recognition system is very essential in terms of surveillance system and many more applications. Nevertheless correctness of these applications still remains the prior challenge. In this paper a simple procedure is proposed to develop a system that monitors a video and detect the objects that are moving and then recognize the objects comparing with our proposed method. Which is primarily targeted to stop stealing of vehicles from parking lot. This particular method can be further extended for many other applications. We use optical flow to detect moving objects and construct a data model of that particular object from different viewpoints. Applying VAIT on those certain regions that has got optical flow we try to figure out the matched object from our data models. Which boosts the performance and reduces the computation significantly. A moving vehicle can be in different angle at a given time. The tangible challenge for this paper is to recognition those objects.
Book Chapters
Service-Learning for Ukraine’s Recovery: Education, Citizenship, and Community Resilience
Practical Cryptography: Algorithms and Implementations Using C++
Datasets
A Novel Dataset for Aspect-based Sentiment Analysis for Teacher Performance Evaluation
TP-NER: A named entity recognition dataset of target and precursor named entities for high-temperature steel
Invited Talk / Workshop
AI Deep Dive: Integrating Artificial Intelligence in Modern Maritime Defense
Delivered a keynote session focused on the strategic application and practical use-cases of Artificial Intelligence (AI) in enhancing naval operations, intelligence gathering, and decision support systems for modern maritime security.
Talk & Guideline Link: View Session Summary
Resources:
Comprehensive Python Programming Workshop: From Fundamentals to Application
Conducted an intensive workshop designed to equip participants with core Python programming skills. Focused heavily on hands-on coding, practical assignments, and applying foundational knowledge to real-world problem domains.
Resources:
Guidelines for the Ethical Use of GenAI for HE and TVET Students in Bangladesh
Organizers: AIUB, JU, BUET, NTU (UK), supported by British Council. Focused on responsible and ethical GenAI integration for higher education and TVET in Bangladesh.
Talk & guideline link: Available here
Resources:
Career in Machine Learning: Opportunities in Bangladesh and the Global Market
Organizers: IEEE CS LU SB Chapter, Leading University, Sylhet
Resources:
First Steps in Research: Research Paper Writing for Beginners
This hands-on workshop covered:
- Selecting a research topic
- Structuring a research paper
- Effective literature review
- Research methodologies
- Citation and referencing
Seeing active student engagement and early research confidence was highly rewarding. The attendees are now better equipped to contribute with well-crafted academic research papers.
Resources (placeholders):














