Optimizing Privacy-Utility Trade-Off in Data Publishing

[thumbnail of Thesis]
Preview
JAHAN, Samsad - THESIS_no signature.pdf - Submitted Version (5MB) | Preview

Jahan, Samsad ORCID logoORCID: https://orcid.org/0000-0001-9921-1630 (2026) Optimizing Privacy-Utility Trade-Off in Data Publishing. PhD thesis, Victoria University.

Abstract

In the contemporary era of big data, digital information is generated, collected, and disseminated at unprecedented scales. The pervasive and often seamless exchange of data across heterogeneous platforms—including, but not limited to, geolocation and health-related information—substantially increases the risk of individual identity disclosure and associated privacy breaches. Such data, therefore, require substantial protection during publication. Researchers have developed various privacy-preserving techniques to protect individual information; however, overly stringent protection often compromises data utility. Most conventional approaches rely on greedy or adaptive strategies to balance privacy and utility, yet these methods exhibit limitations in search capability and adaptability. In this context, three key research problems are identified. First, the inadequate formulation of the privacy–utility trade-off remains a fundamental challenge. Therefore, determining how to better model and optimize the trade-off between privacy and utility for sensitive data constitutes the first research problem. Second, there is a lack of comprehensive Evolutionary Algorithm (EA)-based frameworks for optimizing the privacy–utility trade-off. Hence, determining how to effectively apply EAs to achieve superior privacy–utility solutions for existing Privacy-Preserving Data Publishing (PPDP) problems forms the second research problem. Third, Differential Privacy (DP) has gained significant attention in domains such as Machine Learning (ML) and data mining. However, optimizing privacy budget allocation in DP-based ML models remains an open challenge. Accordingly, determining how to optimally allocate and manage the privacy budget in DP for ML applications represents the third research problem. Based on these research problems, this thesis aims to: i) formulate an improved privacy–utility trade-off model for sensitive data; ii) apply EAs to derive high-quality privacy–utility solutions for PPDP problems; and iii) optimize privacy budget allocation for DP in ML. The overarching objective is to design an efficient privacy-preserving framework that employs EAs and a Multi-Objective Optimization (MOO) approach to generate near-optimal privacy–utility solutions for data publishing. To address the identified research problems and achieve the stated objectives, the k-anonymity and l-diversity standards for medical data were first evaluated. Recommendations were provided to enhance these privacy models through attribute generalization and record suppression. Subsequently, a MOO-based algorithm integrated with a Genetic Algorithm (GA) was developed to generate highquality privacy–utility solutions. Furthermore, a Dynamic Parameter Genetic Algorithm (DPGA) was proposed, integrating the MOO framework with dynamically designed crossover and mutation operators, including a scramble mutation strategy. In addition, an Adaptive Parameter Memetic Algorithm (APMA) was developed, integrating MOO with adaptive crossover and mutation parameters via an adaptive memory-based mechanism, along with a novel local search strategy. These enhancements improve trajectory data privacy–utility performance compared with conventional approaches. Moreover, this study focuses on developing efficient privacy budget allocation strategies for ML applications. To this end, the MOMA-DPK algorithm was proposed, which applies a Memetic Algorithm (MA) within a differentially private K-means clustering framework using a MOO approach to allocate the privacy budget effectively. Additionally, a GA-based privacy budget allocation framework for medical data publishing was designed to optimally distribute the privacy budget across sensitive attributes. EAs have demonstrated strong capability in addressing complex optimization problems, and their integration with MOO offers a promising direction for balancing privacy and utility. The combined adoption of MOO and EAs introduces a relatively underexplored research avenue in PPDP. The proposed methods were evaluated using real-world sensitive datasets and demonstrated comparatively better performance than conventional methods.

Additional Information

Doctor of Philosophy

Item type Thesis (PhD thesis)
URI https://vuir.vu.edu.au/id/eprint/50352
Subjects Current > FOR (2020) Classification > 4602 Artificial intelligence
Current > Division/Research > College of Science and Engineering
Current > Division/Research > Institute for Sustainable Industries and Liveable Cities
Keywords Digital information, privacy, Evolutionary Algorithm, EA, Privacy-Preserving Data Publishing, PPDP
Download/View statistics View download statistics for this item

Search Google Scholar

Repository staff login