GRENZE International Journal of Engineering and Technology
Vol. 12
(2026), Issue 1
Optimizing Frequent Pattern Mining: Dynamic Mapreduce and Pruning Strategies for Apriori Algorithm in Apache Spark and Mapreduce Frameworks
Authors
S Usha Manjari, Vikrant Sabnis, Jay Kumar Jain
Abstract
The exponential growth of big data necessitates efficient frequent pattern mining techniques to uncover valuable insights from large datasets. This research summary synthesizes findings from two studies that enhance the Apriori algorithm's performance in distributed computing environments, specifically Apache Spark and MapReduce frameworks. The first study optimizes the Apriori algorithm by integrating Dynamic MapReduce and pruning techniques within Apache Spark, achieving a significant reduction in execution time to 43.20 seconds with an optimal configuration of 3 mappers and 2 reducers, compared to 83.20 seconds without pruning. The second study introduces AprioriMR, a MapReduce-based approach, demonstrating a dynamic configuration with 5 mappers and 3 reducers yielding an execution time of 83.84 seconds, a marked improvement over the static configuration's 150.37 seconds. Both studies leverage parallel processing and pruning based on the anti-monotone property to reduce computational complexity and enhance scalability. The findings highlight the transformative impact of adaptive resource allocation and pruning strategies in optimizing frequent pattern mining for big data applications, offering practical insights for practitioners in retail, healthcare, and other data-intensive domains. This summary underscores the synergy of distributed computing paradigms and optimization techniques in addressing the scalability challenges of traditional Apriori implementations.
Pages:
1915 - 1920