Loading... Loading...
Grenze Logo
GRENZE International Journal of Engineering and Technology Vol. 12 (2026), Issue 1

Enhancing Frequent Pattern Mining in Big Data using Optimized Apriori Algorithm with Dynamic MapReduce and Pruning Techniques in Apache Spark Environment

Authors

S Usha Manjari, Vikrant Sabnis, Jay Kumar Jain

Abstract

This research focuses on improving the Apriori algorithm for large frequent pattern mining in map reduce big data environment using dynamic map reduce and the pruning techniques on Apache spark. One of the most important algorithms, which cannot be considered a data mining tool without, is the Apriori algorithm used for the identification of frequent item sets and generation of association rules. However, it increases considerably with the number of data that means for large datasets there must be efficient ways of handling these data. The first step of the work is to investigate the influence of changing the number of mappers and reducers in the framework of the offered Dynamic MapReduce Apriori algorithm. In the view of the above variation, the best of the best configuration is determined to be the one that uses 3 mappers and 2 reducers and the corresponding result has an execution time of 83sec. 20 seconds. This supports the need to consider the strategic configuration of Apriori algorithm’s by adjusting expected outcomes from the algorithm. Then, the study incorporates the integration of sophisticated pruning procedures into the Apache Spark milieu to optimize the performance of frequent pattern discovery. With the help of further experiments using similar mapper and reducer con Figurations, it was found out that there were substantial improvements in the setup executed times in all scenarios. Interestingly, when Hadoop MapReduce model is fine-tuned with pruning strategies and Hadoop Map Reduce 3 mappers with 2 reducers model, the execution time is found to be 43. 20 seconds. This significant improvement clearly illustrates the possible positive impacts of pruning in reducing the necessary computation and in improving the algorithms’ performance. The discoveries underscore the possibility of hybridizing Dynamic MapReduce with pruning methods in Apache Spark for enhancing the Apriori algorithm in front ranking in large scale pattern mining on big data. The work provides important findings on how to efficiently perform operations on massive characteristics and frequent itemsets, thus providing solutions for the problem being investigated.