التحسس المضعوط الموجه بالمهام: النظرية والتطبيق
Task Driven Compressive Sensing: Theory & Application
أميمة الدكاك, ديمة شاهين, محي الدين وايناخ
المعهد العالي للعلوم التطبيقية والتكنولوجيا · سوريا
الموضوعات
علوم تطبيقية وتكنولوجية
الملخص
يهدف بحث الدكتوراه "التّحسّس الضغوط الموجه بالمهام: النظرية والتطبيق"إلى دراسة تطبيق التحسّس الضغوط Compressive sensing والترميز المخلخل sparse coding في عدة مهام مختلفة من معالجة إشارة الكلامية، مع استكشاف التكيفات الممكنة لهذه التقنية لكل مهمة بهدف تحسين الأداء. درسنا في هذا البحث تطبيق الترميز المخلخل (والذي هو حالة خاصة من التحسّس الضغوط) في مهمتين من معالجة إشارة الكلامية: (1) تصنيف الصوتيمات باللغة العربية (تعرّف الصوتيمات)، و(2) تعزيز الكلام (إزالة الضجيج). خلال العشر سنوات الماضية، ازدادت الأبحاث المتعلقة بتطبيقات الترميز المخلخل في مجال معالجة الإشارات مثل "تعرّف الوجوه"، وإزالة الضجيج من الصور. يتطلب الترميز المخلخل حل مسألتين جزئيتين: (1) إيجاد مجموعة "كاملة" من الإشارات الأساسية (القاموس) التي تمثل إشارات الاختبار بشكل خلخلي (باستخدام عدد قليل منها)، و(2) إيجاد "الرمز المخلخل" الذي يمثل معاملات التكوين المخلخل (مع العديد من القيم الصفرية) للإشارات الأساسية، بحيث يشكل هذا التكوين أقرب تقريب للإشارة المستهدفة بأصغر خطأ ممكن. قدمنا في هذا البحث مساهمتين جديدتين في مجال تطبيق الترميز المخلخل في معالجة إشارة الكلامية: أولاً، لتصنيف الصوتيمات، اقترحنا نظام تصنيف جديد يعتمد على الترميز المخلخل كميزات تمييزية وSparse Representation Classifier (SRC). يتكوّن التصنيف من مرحلتين: (1) مرحلة تدريب حيث يتم بناء/تعلم القاموس، و(2) مرحلة تعرّف حيث يتم حساب الترميز المخلخل لإشارة الصوتيم الاختبارية على القاموس المختار، ثم يُستخدم لتحديد الفئة الصوتيمية لإشارة الاختبار. في مرحلة التدريب، درسنا ثلاثة أنواع مختلفة من القواميس: (a) قاموس الأمثلة exemplars dictionary من متجهات ميزات الصوتيمات التدريبية، (b) قاموس K-Singular Value Decomposition (K-SVD)، و(c) قاموس Fisher Discriminative Dictionary Learning (FDDL). لتقييم أداء نظام تصنيف الصوتيمات المقترح، استخدمنا مجموعتين من بيانات الصوتيمات العربية، تحتويان على 33 صوتيمًا. أظهرت التجارب أن "قاموس الأمثلة" يعطي أفضل أداء من حيث خطأ التصنيف، بينما القواميس المتعلمة تتميز بزمن تصنيف أقل. نظام التصنيف المقترح يتفوق على أداء نظام الشبكات العصبية Echo State Neural Networks الذي تم تطبيقه على نفس مجموعتي بيانات الصوتيم العربية. ثانيًا، لتعزيز الكلام، اقترحنا خوارزمية جديدة لتعلم القاموس Incoherent Discriminative Dictionary Learning (IDDL) لنمذجة كل من الكلام والضجيج. صياغة مسألة تعلم القاموس المقترحة تحتوي على دالة تكلفة تأخذ بعين الاعتبار كلًا من أخطاء "التشويش على المصدر" و"تشويه المصدر"، مع مصطلح تنظيم يعاقب التماسك بين القواميس الفرعية للكلام والضجيج. استخدمنا خوارزمية IDDL المقترحة في نظام تعزيز الكلام المشرف الذي يحتوي على مرحلتين: (1) مرحلة تدريب حيث نتعلم قاموس IDDL المكوّن من قاموسين فرعيين لنمذجة طيف السعة لكل من الكلام النظيف والضجيج، و(2) مرحلة تعزيز حيث يتم حساب الرموز المخلخلة لطيف السعة للإشارة الملوثة على القاموس المتعلم، ثم تُضرب الرموز المخلخلة بالقواميس الفرعية لإيجاد تقدير أولي لكل من طيف سعة الكلام النظيف والضجيج. أخيرًا، يُستخدم مرشح Wiener لتنقية تقدير الكلام النظيف. أظهرت التجارب على مجموعة بيانات NOIZUES (التي تحتوي على إشارات كلامية ملوثة بثمانية أنواع مختلفة من الضجيج)، باستخدام مقياسين موضوعيين لتعزيز الكلام: frequency-weighted segmental SNR (FwSegSNR) وPerceptual Evaluation of Speech Quality (PESQ)، أن خوارزمية IDDL المقترحة تتفوق على خوارزميات تعلم القاموس الأخرى المختبرة (K-SVD، GDL، وFDDL) في معظم حالات أنواع الضجيج المدروسة
روابط وملفات
التعريف والنوع
- رقم الوثيقة
- fd601da7-e2b8-4ca4-8cd5-e447b6f35f21
- رقم العقد
- 0
- نوع الوسائط
- Crawler
- نوع المحتوى
- الرسائل العلمية
- صيغة المصدر
- رسائل دكتوراة
- نوع الملف
- pdf text
- أسماء الملفات
- 841141_1.pdf
بيانات النشر
- ترجمة العنوان
- Task Driven Compressive Sensing: Theory & Application
- ألقاب المؤلفين
- [{"name_ar":"أميمة الدكاك","title_ar":"اشراف","title_en":"Supervision"},{"name_ar":"ديمة شاهين","title_ar":"اعداد","title_en":"Preparation"},{"name_ar":"محي الدين وايناخ","title_ar":"اشراف","title_en":"Supervision"}]
- اللغة
- Arabic
المصدر والدورية
- اسم المصدر
- التحسس المضعوط الموجه بالمهام: النظرية والتطبيق
المحتوى والصفحات
- عدد الصفحات
- 0
- ترجمة الملخص
- The PhD research "Compressive Sensing: Practical Application" studies the application of Compressive Sensing and Sparse Coding in various speech signal processing tasks, exploring possible adaptations of this technology for each task to optimize performance. In this research, we studied the application of sparse coding (a special case of compressive sensing) in two speech signal processing tasks: Arabic phoneme classification (phoneme recognition). Speech enhancement (denoising). Over the past ten years, research on sparse coding applications in signal processing, such as face recognition and image denoising, has increased. Sparse coding requires solving two sub-problems: Finding a large set of “basic signals” (dictionary) that can represent the test signals sparsely (using only a few of them). Finding the "sparse code" representing the coefficients of the linear combination of the basic signals (with many zeros), approximating the target signal as closely as possible. This research contributes in two ways to the application of sparse coding in speech signal processing: First, for phoneme classification: We proposed a new classification system based on sparse coding as discriminative features and Sparse Representation Classifier (SRC). The classification is composed of two stages: Training stage where the dictionary is built/learned. Recognition stage where the sparse code of the test phoneme signal is calculated on the chosen dictionary, then used to decide which phoneme class the test signal belongs to. In the training stage, we considered three types of dictionaries: a) Exemplars dictionary from training phoneme feature vectors. b) K-Singular Value Decomposition (K-SVD) dictionary. c) Fisher Discriminative Dictionary Learning (FDDL) dictionary. To evaluate the performance of the proposed phoneme classification system, we used two Arabic phoneme datasets containing 33 phonemes. Experiments show that the exemplars dictionary gives the best performance in terms of classification error, while the learned dictionaries have shorter classification time. The proposed classification system outperforms the Echo State Neural Networks applied to the same two Arabic phoneme datasets. Second, for speech enhancement: We proposed a new dictionary learning algorithm, Incoherent Discriminative Dictionary Learning (IDDL), to model both speech and noise. The proposed dictionary learning formulation includes a cost function that accounts for both "source confusion" and "source distortion" errors, with a regularization term penalizing coherence between the speech and noise sub-dictionaries. We used the IDDL algorithm in a supervised speech enhancement system consisting of two stages: Training stage to learn the IDDL dictionary composed of two sub-dictionaries modeling the amplitude spectrum of both clean speech and noise. Enhancement stage where the sparse codes of the amplitude spectrum of the noisy signal are calculated on the learned dictionary, then multiplied by the sub-dictionaries to obtain an initial estimate of both clean speech and noise amplitude spectrum. Finally, a Wiener filter is used to refine the clean speech estimate. Experiments on the NOIZUES dataset (containing speech signals contaminated with eight different types of noise), using two objective speech enhancement measures: frequency-weighted segmental SNR (FwSegSNR) and Perceptual Evaluation of Speech Quality (PESQ), demonstrate that the proposed IDDL algorithm outperforms other tested dictionary learning algorithms (K-SVD, GDL, and FDDL) in many cases.
- كلمات الباحثين
- التحسس الضغوط، التمييز الخلخل، خوارزميات تعلم القاموس، تعرّف الصوتيمات، برستُ الكلاّ، تحسين الكلام، تصنيف الصوتيمات، الترميز المتناثر، تعلم قاموس Fisher، K-SVD، Incoherent Discriminative Dictionary Learning.
إشراف وإعداد
- الإشراف
- أميمة الدكاك, محي الدين وايناخ
- الإعداد
- ديمة شاهين
الاقتباسات الببليوغرافية
APA
MLA