| Supervised Speech Separation Based on Deep Learning: An Overview |
101 |
| Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation |
61 |
| End-to-End Waveform Utterance Enhancement for Direct Evaluation Metrics Optimization by Fully Convolutional Neural Networks |
34 |
| Detection and Classification of Acoustic Scenes and Events: Outcome of the DCASE 2016 Challenge |
26 |
| A New Framework for CNN-Based Speech Enhancement in the Time Domain |
21 |
| Sound Event Detection and Time-Frequency Segmentation from Weakly Labelled Data |
19 |
| Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks |
18 |
| Gated Residual Networks With Dilated Convolutions for Monaural Speech Enhancement |
17 |
| Speech Emotion Classification Using Attention-Based LSTM |
17 |
| Adaptive Pooling Operators for Weakly Labeled Sound Event Detection |
17 |
| Divide and Conquer: A Deep CASA Approach to Talker-Independent Monaural Speaker Separation |
13 |
| Two-Stage Deep Learning for Noisy-Reverberant Speech Enhancement |
13 |
| Phase-Aware Speech Enhancement Based on Deep Neural Networks |
13 |
| Active Noise Control Over Space: A Wave Domain Approach |
12 |
| Combining Spectral and Spatial Features for Deep Learning Based Blind Speaker Separation |
12 |
| Optimization of RNN-Based Speech Activity Detection |
12 |
| A Hierarchy-to-Sequence Attentional Neural Machine Translation Model |
11 |
| Refining Word Embeddings Using Intensity Scores for Sentiment Analysis |
11 |
| Semisupervised Autoencoders for Speech Emotion Recognition |
11 |
| Text-Independent Speaker Verification Based on Triplet Convolutional Neural Network Embeddings |
11 |
| A Multiobjective Learning and Ensembling Approach to High-Performance Speech Enhancement With Compact Neural Network Architectures |
11 |
| Domain Adversarial for Acoustic Emotion Recognition |
11 |
| Waveform Modeling and Generation Using Hierarchical Recurrent Neural Networks for Speech Bandwidth Extension |
10 |
| Acoustic SLAM |
10 |
| Evaluation and Comparison of Late Reverberation Power Spectral Density Estimators |
10 |
| An Overview of Lead and Accompaniment Separation in Music |
10 |
| Adaptive Very Deep Convolutional Residual Network for Noise Robust Speech Recognition |
10 |
| DNN-Based Source Enhancement to Increase Objective Sound Quality Assessment Score |
10 |
| A Low-Complexity Robust Beamforming Using Diagonal Unloading for Acoustic Source Localization |
9 |
| Linear System Identification Based on a Kronecker Product Decomposition |
9 |
| Robust Binaural Localization of a Target Sound Source by Combining Spectral Source Models and Deep Neural Networks |
9 |
| Spread Spectrum Audio Watermarking Using Multiple Orthogonal PN Sequences and Variable Embedding Strengths and Polarities |
9 |
| Speech Enhancement Based on Teacher-Student Deep Learning Using Improved Speech Presence Probability for Noise-Robust Speech Recognition |
9 |
| Weakly Labelled AudioSet Tagging With Attention Neural Networks |
9 |
| Insights Into Frequency-Invariant Beamforming With Concentric Circular Microphone Arrays |
8 |
| A Neural Approach to Source Dependence Based Context Model for Statistical Machine Translation |
8 |
| Robust Speaker Localization Guided by Deep Learning-Based Time-Frequency Masking |
8 |
| Sentiment Lexicon Construction With Hierarchical Supervision Topic Model |
8 |
| Sequence-to-Sequence Acoustic Modeling for Voice Conversion |
8 |
| Analysis of the Reconstruction of Sparse Signals in the DCT Domain Applied to Audio Signals |
8 |
| Deep Learning for Talker-Dependent Reverberant Speaker Separation: An Empirical Study |
8 |
| Sentence Selection and Weighting for Neural Machine Translation Domain Adaptation |
8 |
| Recursive Least-Squares Algorithms for the Identification of Low-Rank Systems |
8 |
| Sample Efficient Deep Reinforcement Learning for Dialogue Systems With Large Action Spaces |
8 |
| Unsupervised Speech Representation Learning Using WaveNet Autoencoders |
8 |
| TDOA-Based Multiple Acoustic Source Localization Without Association Ambiguity |
7 |
| Sound Event Recognition Using Auditory-Receptive-Field Binary Pattern and Hierarchical-Diving Deep Belief Network |
7 |
| Curriculum Learning for Speech Emotion Recognition From Crowdsourced Labels |
7 |
| Improving Aspect Term Extraction With Bidirectional Dependency Tree Representation |
7 |
| Incorporating Statistical Machine Translation Word Knowledge Into Neural Machine Translation |
7 |