Schoobrary رجوع
العودة إلى البحث
ماستر (LMD) الانجليزية 2021 011c76da-5344-4b44-98a1-71dc71841c0c

Privacy-Preserving Documents Matching Based on Hybrid Methods

Ayad Ibrahim Abdulsada, Duaa Fadhl Najm

كلية التربية للعلوم الصرفة-جامعة البصرة · العراق

الموضوعات

علوم تطبيقية وتكنولوجية

الملخص

Document similarity detection is an appealing method used for many real-world application domains such as intellectual property protection, digital file systems, article plagiarism, and digital shopping. Usually, most of the current methods work on non-private documents, which are suitable for only environments where documents are publicly accessed. Unfortunately, this assumption is not suitable for many real-world applications, e.g., document similarity detection between two security agencies, where documents are supposed to be confidential. Inthe last few years, cryptographers developed privacy-preserving protocols for detecting similar documents without compromising privacy. However, existing protocols are either not secure or impractical. In this thesis, the work focuses on the issue of identifying the resemblance between textual documents belonging to two parties, who are unwilling to disclose their private content while maintaining efficiency. Three constructions have been presented. All of these constructions employ the N-gram model to represent textual documents as sets. The first construction used large N-gramsets and the secondconstruction is reduced the large N-gram sets by utilizing the Min-hash technique. A twoprotocol for secure intersection set (PSI-CA andpaillier based intersection) is used to find the common elements between two sets. Such information is used to estimate the similarity scores. The third construction enhances security by hiding the actual values of the similarity scores.Several experimentshave been conducted on real collection to illustrate the performance of the proposed constructions. Experimental results show that the PSI-CA is faster than paillier based intersection. The Document similarity detection is an appealing method used for many real-world application domains such as intellectual property protection, digital file systems, article plagiarism, and digital shopping. Usually, most of the current methods work on non-private documents, which are suitable for only environments where documents are publicly accessed. Unfortunately, this assumption is not suitable for many real-world applications, e.g., document similarity detection between two security agencies, where documents are supposed to be confidential. Inthe last few years, cryptographers developed privacy-preserving protocols for detecting similar documents without compromising privacy. However, existing protocols are either not secure or impractical. In this thesis, the work focuses on the issue of identifying the resemblance between textual documents belonging to two parties, who are unwilling to disclose their private content while maintaining efficiency. Three constructions have been presented. All of these constructions employ the N-gram model to represent textual documents as sets. The first construction used large N-gramsets and the secondconstruction is reduced the large N-gram sets by utilizing the Min-hash technique. A twoprotocol for secure intersection set (PSI-CA andpaillier based intersection) is used to find the common elements between two sets. Such information is used to estimate the similarity scores. The third construction enhances security by hiding the actual values of the similarity scores.Several experimentshave been conducted on real collection to illustrate the performance of the proposed constructions. Experimental results show that the PSI-CA is faster than paillier based intersection. The Document similarity detection is an appealing method used for many real-world application domains such as intellectual property protection, digital file systems, article plagiarism, and digital shopping. Usually, most of the current methods work on non-private documents, which are suitable for only environments where documents are publicly accessed. Unfortunately, this assumption is not suitable for many real-world applications, e.g., document similarity detection between two security agencies, where documents are supposed to be confidential. Inthe last few years, cryptographers developed privacy-preserving protocols for detecting similar documents without compromising privacy. However, existing protocols are either not secure or impractical. In this thesis, the work focuses on the issue of identifying the resemblance between textual documents belonging to two parties, who are unwilling to disclose their private content while maintaining efficiency. Three constructions have been presented. All of these constructions employ the N-gram model to represent textual documents as sets. The first construction used large N-gramsets and the secondconstruction is reduced the large N-gram sets by utilizing the Min-hash technique. A twoprotocol for secure intersection set (PSI-CA andpaillier based intersection) is used to find the common elements between two sets. Such information is used to estimate the similarity scores. The third construction enhances security by hiding the actual values of the similarity scores.Several experimentshave been conducted on real collection to illustrate the performance of the proposed constructions. Experimental results show that the PSI-CA is faster than paillier based intersection. The first construction has the best overall effectiveness because it employs the 3-gram set to represent each document. The third construction has minimum effectiveness since it employs only the intersection operation to measure the similarity but it's a high level of security because it is hiding the actual values of the similarity scores. The second construction was faster than other constructions in estimating approximate similarity over the resulting sets with minimum computation and communication costs.

التعريف والنوع

رقم الوثيقة
011c76da-5344-4b44-98a1-71dc71841c0c
رقم العقد
0
نوع الوسائط
Crawler
نوع المحتوى
الرسائل العلمية
صيغة المصدر
ماستر (LMD)
نوع الملف
pdf image
أسماء الملفات
011c76da-5344-4b44-98a1-71dc71841c0c_1.pdf

بيانات النشر

ألقاب المؤلفين
[{"name_ar":"Ayad Ibrahim Abdulsada","title_ar":"اشراف","title_en":"Supervision"},{"name_ar":"Duaa Fadhl Najm ","title_ar":"اعداد","title_en":"Preparation"}]
اللغة
English

المصدر والدورية

اسم المصدر
Privacy-Preserving Documents Matching Based on Hybrid Methods

المحتوى والصفحات

عدد الصفحات
0

إشراف وإعداد

الإشراف
Ayad Ibrahim Abdulsada
الإعداد
Duaa Fadhl Najm

الاقتباسات الببليوغرافية

APA

Ayad Ibrahim Abdulsada و Duaa Fadhl Najm . (2021). Privacy-Preserving Documents Matching Based on Hybrid Methods . أطروحة(ماستر (LMD)). كلية التربية للعلوم الصرفة-جامعة البصرة. العراق.

MLA

Ayad Ibrahim Abdulsada و Duaa Fadhl Najm . Privacy-Preserving Documents Matching Based on Hybrid Methods . 2021. كلية التربية للعلوم الصرفة-جامعة البصرة، ماستر (LMD).