Privacy-Preserving Documents Matching Based on Hybrid Methods
Ayad Ibrahim Abdulsada, Duaa Fadhl Najm
كلية التربية للعلوم الصرفة-جامعة البصرة · العراق
الموضوعات
علوم تطبيقية وتكنولوجية
الملخص
Document similarity detection is an appealing method used for many real-world application domains such as intellectual property protection, digital file systems, article plagiarism, and digital shopping. Usually, most of the current methods work on non-private documents, which are suitable for only environments where documents are publicly accessed. Unfortunately, this assumption is not suitable for many real-world applications, e.g., document similarity detection between two security agencies, where documents are supposed to be confidential. Inthe last few years, cryptographers developed privacy-preserving protocols for detecting similar documents without compromising privacy. However, existing protocols are either not secure or impractical. In this thesis, the work focuses on the issue of identifying the resemblance between textual documents belonging to two parties, who are unwilling to disclose their private content while maintaining efficiency. Three constructions have been presented. All of these constructions employ the N-gram model to represent textual documents as sets. The first construction used large N-gramsets and the secondconstruction is reduced the large N-gram sets by utilizing the Min-hash technique. A twoprotocol for secure intersection set (PSI-CA andpaillier based intersection) is used to find the common elements between two sets. Such information is used to estimate the similarity scores. The third construction enhances security by hiding the actual values of the similarity scores.Several experimentshave been conducted on real collection to illustrate the performance of the proposed constructions. Experimental results show that the PSI-CA is faster than paillier based intersection. The Document similarity detection is an appealing method used for many real-world application domains such as intellectual property protection, digital file systems, article plagiarism, and digital shopping. Usually, most of the current methods work on non-private documents, which are suitable for only environments where documents are publicly accessed. Unfortunately, this assumption is not suitable for many real-world applications, e.g., document similarity detection between two security agencies, where documents are supposed to be confidential. Inthe last few years, cryptographers developed privacy-preserving protocols for detecting similar documents without compromising privacy. However, existing protocols are either not secure or impractical. In this thesis, the work focuses on the issue of identifying the resemblance between textual documents belonging to two parties, who are unwilling to disclose their private content while maintaining efficiency. Three constructions have been presented. All of these constructions employ the N-gram model to represent textual documents as sets. The first construction used large N-gramsets and the secondconstruction is reduced the large N-gram sets by utilizing the Min-hash technique. A twoprotocol for secure intersection set (PSI-CA andpaillier based intersection) is used to find the common elements between two sets. Such information is used to estimate the similarity scores. The third construction enhances security by hiding the actual values of the similarity scores.Several experimentshave been conducted on real collection to illustrate the performance of the proposed constructions. Experimental results show that the PSI-CA is faster than paillier based intersection. The Document similarity detection is an appealing method used for many real-world application domains such as intellectual property protection, digital file systems, article plagiarism, and digital shopping. Usually, most of the current methods work on non-private documents, which are suitable for only environments where documents are publicly accessed. Unfortunately, this assumption is not suitable for many real-world applications, e.g., document similarity detection between two security agencies, where documents are supposed to be confidential. Inthe last few years, cryptographers developed privacy-preserving protocols for detecting similar documents without compromising privacy. However, existing protocols are either not secure or impractical. In this thesis, the work focuses on the issue of identifying the resemblance between textual documents belonging to two parties, who are unwilling to disclose their private content while maintaining efficiency. Three constructions have been presented. All of these constructions employ the N-gram model to represent textual documents as sets. The first construction used large N-gramsets and the secondconstruction is reduced the large N-gram sets by utilizing the Min-hash technique. A twoprotocol for secure intersection set (PSI-CA andpaillier based intersection) is used to find the common elements between two sets. Such information is used to estimate the similarity scores. The third construction enhances security by hiding the actual values of the similarity scores.Several experimentshave been conducted on real collection to illustrate the performance of the proposed constructions. Experimental results show that the PSI-CA is faster than paillier based intersection. The first construction has the best overall effectiveness because it employs the 3-gram set to represent each document. The third construction has minimum effectiveness since it employs only the intersection operation to measure the similarity but it's a high level of security because it is hiding the actual values of the similarity scores. The second construction was faster than other constructions in estimating approximate similarity over the resulting sets with minimum computation and communication costs.
روابط وملفات
التعريف والنوع
- رقم الوثيقة
- 011c76da-5344-4b44-98a1-71dc71841c0c
- رقم العقد
- 0
- نوع الوسائط
- Crawler
- نوع المحتوى
- الرسائل العلمية
- صيغة المصدر
- ماستر (LMD)
- نوع الملف
- pdf image
- أسماء الملفات
- 011c76da-5344-4b44-98a1-71dc71841c0c_1.pdf
بيانات النشر
- ألقاب المؤلفين
- [{"name_ar":"Ayad Ibrahim Abdulsada","title_ar":"اشراف","title_en":"Supervision"},{"name_ar":"Duaa Fadhl Najm ","title_ar":"اعداد","title_en":"Preparation"}]
- اللغة
- English
المصدر والدورية
- اسم المصدر
- Privacy-Preserving Documents Matching Based on Hybrid Methods
المحتوى والصفحات
- عدد الصفحات
- 0
إشراف وإعداد
- الإشراف
- Ayad Ibrahim Abdulsada
- الإعداد
- Duaa Fadhl Najm
الاقتباسات الببليوغرافية
APA
MLA