[1] Goodfellow, I., Bengio, Y., and Courville, A., Deep learning, MIT press, 2016.
[2] Wang, Y., Yao, Q., Kwok, J., and Ni, L. M., "Generalizing from a Few Examples: A Survey on Few-Shot Learning",
ACM Computing Surveys (CSUR), Vol. 53, No. 3, Apr. 2019, Accessed: Jan. 23, 2022. [Online]. Available:
http://arxiv.org/abs/1904.05046.
[3] Lake, B. M., Salakhutdinov, R., and Tenenbaum, J. B., "Human-level concept learning through probabilistic program induction", Science, Vol. 350, No. 6266, pp. 1332–8, doi: 10.1126/science.aab3050, Dec, 2015.
[4] Jiang, X., Ding, L., Havaei, M., Jesson, A., and Matwin, S., "Task Adaptive Metric Space for Medium-Shot Medical Image Classification", in MICCAI, pp. 147–155. doi: 10.1007/978-3-030-32239-7_17, 2019.
[5] Li, X., Sun, Z., Xue, J.-H., and Ma, Z., "A concise review of recent few-shot meta-learning methods", Neurocomputing, Vol. 456, pp. 463–468, doi: 10.1016/j.neucom.2020.05.114, Oct, 2021.
[6] Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K., and Müller, K.-R., Explainable AI: interpreting, explaining and visualizing deep learning, Vol. 11700. Springer Nature, 2019.
[7] Zhang, Y., Tino, P., Leonardis, A., and Tang, K., "A Survey on Neural Network Interpretability", IEEE Transactions on Emerging Topics in Computational Intelligence, Vol. 5, No. 5, pp. 726–742, doi: 10.1109/TETCI.2021.3100641, Oct, 2021.
[8] Patacchiola, M., Turner, J., Crowley, E. J., O’Boyle, M., and Storkey, A., "Bayesian Meta-Learning for the Few-Shot Setting via Deep Kernels",
34th Conference on Neural Information Processing Systems (NeurIPS 2020), Oct. 2019, Accessed: Feb. 04, 2021. [Online]. Available:
http://arxiv.org/abs/1910.05199.
[9] Huisman, M., van Rijn, J. N., and Plaat, A., "A survey of deep meta-learning", Artificial Intelligence Review, Vol. 54, No. 6, pp. 4483–4541, doi: 10.1007/s10462-021-10004-4, Aug. 2021.
[10] Santoro, A., Bartunov, S., Botvinick, M., Wierstra, D., and Lillicrap, T., "One-shot Learning with Memory-Augmented Neural Networks",
33rd International Conference on Machine Learning, ICML 2016, Vol. 4, pp. 2740–2751, May 2016, Accessed: Oct. 16, 2019. [Online]. Available:
http://arxiv.org/abs/1605.06065
[11] Munkhdalai, T., and Yu, H., "Meta Networks",
34th International Conference on Machine Learning, ICML, Vol. 5, pp. 3933–3943, Mar. 2017, Accessed: Oct. 18, 2019. [Online]. Available:
http://arxiv.org/abs/1703.00837
[12] Andrychowicz, M.,
et al., "Learning to learn by gradient descent by gradient descent",
Advances in Neural Information Processing Systems, pp. 3988–3996, Jun. 2016, Accessed: Oct. 16, 2019. [Online]. Available:
http://arxiv.org/abs/1606.04474
[13] Ravi, S., and Larochelle, H., "Optimization as a model for few-shot learning", 2017.
[14] Vinyals, O., Blundell, C., Lillicrap, T., Kavukcuoglu, K., and Wierstra, D., "Matching Networks for One Shot Learning",
Advances in Neural Information Processing Systems, pp. 3637–3645, Jun. 2016, Accessed: Oct. 18, 2019. [Online]. Available:
http://arxiv.org/abs/1606.04080
[15] Snell, J., Swersky, K., and Zemel, R. S., "Prototypical Networks for Few-shot Learning",
Advances in Neural Information Processing Systems, Vol. 2017-Decem, pp. 4078–4088, Mar. 2017, Accessed: Oct. 18, 2019. [Online]. Available:
http://arxiv.org/abs/1703.05175
[16] Sung, F., Yang, Y., Zhang, L., Xiang, T., Torr, P. H. S., and Hospedales, T. M., "Learning to Compare: Relation Network for Few-Shot Learning", in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1199–1208. doi: 10.1109/CVPR.2018.00131, Jun, 2018.
[17] Gordon, J., Bronskill, J., Bauer, M., Nowozin, S., and Turner, R. E., "Meta-Learning Probabilistic Inference For Prediction",
7th In International Conference on Learning Representations, ICLR, May 2018, Accessed: Jan. 23, 2022. [Online]. Available:
http://arxiv.org/abs/1805.09921
[18] Finn, C., Xu, K., and Levine, S., "Probabilistic Model-Agnostic Meta-Learning", Advances in Neural Information Processing Systems, pp. 9516–9527, Jun. 2018.
[19] Garnelo, M.,
et al., "Conditional Neural Processes",
35th International Conference on Machine Learning, ICML 2018, Vol. 4, pp. 2738–2747, Jul. 2018, Accessed: Apr. 21, 2020. [Online]. Available:
http://arxiv.org/abs/1807.01613
[20] Edwards, H., and Storkey, A., "Towards a Neural Statistician",
5th International Conference on Learning Representations, ICLR, Accessed: Apr. 26, 2020, Jun, 2016. [Online]. Available:
http://arxiv.org/abs/1606.02185
[21] Finn, C., Abbeel, P., and Levine, S., "Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks",
34th International Conference on Machine Learning, ICML 2017, Vol. 3, pp. 1856–1868, Mar. 2017, Accessed: Oct. 19, 2019. [Online]. Available:
http://arxiv.org/abs/1703.03400
[22] Kim, T., Yoon, J., Dia, O., Kim, S., Bengio, Y., and Ahn, S., "Bayesian Model-Agnostic Meta-Learning", Advances in Neural Information Processing Systems, pp. 7332–7342, Jun. 2018.
[23] Grant, E., Finn, C., Levine, S., Darrell, T., and Griffiths, T., "Recasting Gradient-Based Meta-Learning as Hierarchical Bayes", 6th International Conference on Learning Representations, ICLR, Jan. 2018.
[24] Ravi, S., and Beatson, A., "Amortized bayesian meta-learning", 2018.
[25] Raghu, A., Raghu, M., Bengio, S., and Vinyals, O., "Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML", 8th International Conference on Learning Representations, ICLR 2020, Sep. 2019.
[26] Oh, J., Yoo, H., Kim, C., and Yun, S., "BOIL: Towards Representation Change for Few-shot Learning", 2021.
[27] Zintgraf, L., Shiarlis, K., Kurin, V., Hofmann, K., and Whiteson, S., "Fast context adaptation via meta-learning", in 36th International Conference on Machine Learning, ICML 2019, Vol. 2019-June, pp. 13262–13276, 2019.
[28] Nichol, A., Achiam, J., and Schulman, J., "On First-Order Meta-Learning Algorithms", Mar. 2018, Accessed: Apr. 13, 2020. [Online]. Available:
http://arxiv.org/abs/1803.02999
[29] Rajeswaran, A., Finn, C., Kakade, S., and Levine, S., "Meta-Learning with Implicit Gradients",
Advances in Neural Information Processing Systems 32(pp. 113-124), Sep. 2019, Accessed: Aug. 16, 2020. [Online]. Available:
http://arxiv.org/abs/1909.04630
[30] Lee, K., Maji, S., Ravichandran, A., and Soatto, S., "Meta-Learning With Differentiable Convex Optimization", in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10649–10657. doi: 10.1109/CVPR.2019.01091, Jun, 2019.
[31] Bertinetto, L., Henriques, J. F., Torr, P. H. S., and Vedaldi, A., "Meta-learning with differentiable closed-form solvers",
7th In International Conference on Learning Representations, ICLR, May 2018, [Online]. Available:
http://arxiv.org/abs/1805.08136
[32] Gai, S., and Wang, D., "Sparse Model-Agnostic Meta-Learning Algorithm for Few-Shot Learning", in 2019 2nd China Symposium on Cognitive Computing and Hybrid Intelligence (CCHI), pp. 127–130, Sep. 2019. doi: 10.1109/CCHI.2019.8901909.
[33] Madan, A., and Prasad, R., "B-Small: A Bayesian Neural Network Approach to Sparse Model-Agnostic Meta-Learning", ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 2730–2734, 2021.
[34] Tian, H., Liu, B., Yuan, X. -T., and Liu, Q., "Meta-learning with Network Pruning", in Computer Vision – ECCV 2020. Lecture Notes in Computer Science, Springer, Cham, pp. 675–700. doi: 10.1007/978-3-030-58529-7_40, 2020.
[35] Sabzevar, R. Z., Ghiasi-Shirazi, K., and Harati, A., "Prototype-based interpretation of the functionality of neurons in winner-take-all neural networks",
ArXiv, vol. abs/2008.08750, Aug. 2020, [Online]. Available:
http://arxiv.org/abs/2008.08750
[36] Chen, C., Li, O., Tao, C., Barnett, A. J., Su, J., and Rudin, C., "This Looks Like That: Deep Learning for Interpretable Image Recognition",
Advances in Neural Information Processing Systems (NeurIPS 2018)), Vol. 32, pp. 8930–8941, Jun. 2019, [Online]. Available:
http://arxiv.org/abs/1806.10574
[37] Cao, K., Brbic, M., and Leskovec, J., "Concept Learners for Few-Shot Learning", International Conference on Learning Representations (ICLR), 2021.
[38] Koh, P. W., and Liang, P., "Understanding Black-box Predictions via Influence Functions",
International Conference on Machine Learning, pp. 1885–1894, Mar. 2017, [Online]. Available:
http://arxiv.org/abs/1703.04730
[39] Yeh, C. -K., Kim, J. S., Yen, I. E. H., and Ravikumar, P., "Representer Point Selection for Explaining Deep Neural Networks",
Advances in Neural Information Processing Systems, Vol. 31, Nov. 2018, [Online]. Available:
http://arxiv.org/abs/1811.09720
[40] Arik, S. O., and Pfister, T., "ProtoAttend: Attention-Based Prototypical Learning",
Journal of Machine Learning Research, Vol. 21, pp. 1–35, Feb. 2020, [Online]. Available:
http://arxiv.org/abs/1902.06292
[41] Vaswani, A., et al., "Attention is all you need", in Advances in neural information processing systems, pp. 5998–6008, 2017.
[42] Tsai, Y. -H. H., Bai, S., Yamada, M., Morency, L. -P., and Salakhutdinov, R. "Transformer Dissection: An Unified Understanding for Transformer’s Attention via the Lens of Kernel", Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2019.
[43] Chen, Y., Zeng, Q., Ji, H., and Yang, Y., "Skyformer: Remodel Self-Attention with Gaussian Kernel and Nystrom Method", Advances in Neural Information Processing Systems, Vol. 34, 2021.
[44] Song, K., Jung, Y., Kim, D., and Moon, I. -C., "Implicit kernel attention", in Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35, No. 11, pp. 9713–9721, 2021.
[45] Choromanski, K. M., et al., "Rethinking Attention with Performers", International Conference on Learning Representations, 2021.
[46] Schlkopf, B., Smola, A. J., and Bach, F., Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond. The MIT Press, 2018.
[47] Wilson, A. G., Hu, Z., Salakhutdinov, R., and Xing, E. P., "Deep Kernel Learning",
Artificial intelligence and statistics, AISTATS, pp. 370–378, Nov. 2016, [Online]. Available:
http://arxiv.org/abs/1511.02222
[48] Tossou, P., Dura, B., Laviolette, F., Marchand, M., and Lacoste, A., "Adaptive deep kernel learning", arXiv preprint arXiv:1905.12131, 2019.
[49] Salakhutdinov, R., and Hinton, G. E., "Using Deep Belief Nets to Learn Covariance Kernels for Gaussian Processes", in NIPS, Vol. 7, pp. 1249–1256, 2007.
[50] Calandra, R., Peters, J., Rasmussen, C. E., and Deisenroth, M. P., "Manifold Gaussian processes for regression", in 2016 International Joint Conference on Neural Networks (IJCNN), pp. 3338–3345, 2016.
[51] Cortes, C., and Vapnik, V., "Support-vector networks", Machine Learning, Vol. 20, No. 3, pp. 273–297, doi: 10.1007/bf00994018, Sep. 1995.
[52] Rasmussen, C. E., and Williams, C. K. I.,
Gaussian Processes for Machine Learning. The MIT Press, 2006. [Online]. Available:
http://www.gaussianprocess.org/gpml/
[53] Quinonero-Candela, J., and Rasmussen, C. E., "A unifying view of sparse approximate Gaussian process regression", The Journal of Machine Learning Research, Vol. 6, pp. 1939–1959, 2005.
[54] Liu, H., Ong, Y. -S., Shen, X., and Cai, J., "When Gaussian process meets big data: A review of scalable GPs", IEEE Trans Neural Netw Learn Syst, Vol. 31, No. 11, pp. 4405–4423, 2020.
[55] Smola, A., and Bartlett, P., "Sparse greedy Gaussian process regression", Adv Neural Inf Process Syst, Vol. 13, 2000.
[56] Seeger, M. W., Williams, C. K. I., and Lawrence, N. D., "Fast forward selection to speed up sparse Gaussian process regression", in International Workshop on Artificial Intelligence and Statistics, pp. 254–261, 2003.
[57] Keerthi, S. S., and Chu, W., "A matching pursuit approach to sparse Gaussian process regression", Adv Neural Inf Process Syst, Vol. 18, 2005.
[58] Snelson, E., and Ghahramani, Z., "Sparse Gaussian processes using pseudo-inputs", Adv Neural Inf Process Syst, Vol. 18, 2005.
[59] Williams, C., and Seeger, M., "Using the Nyström Method to Speed Up Kernel Machines", in Advances in Neural Information Processing Systems, Vol. 13, pp. 682–688, 2000.
[60] Tipping, M. E., "Sparse Bayesian Learning and the Relevance Vector Machine", J. Mach. Learn. Res., Vol. 1, pp. 211–244, 2001.
[61] Bishop, C. M., Pattern Recognition and Machine Learning (Information Science and Statistics). Berlin, Heidelberg: Springer-Verlag, 2006.
[62] Tipping, M. E., and Faul, A. C., "Fast marginal likelihood maximisation for sparse Bayesian models", International workshop on artificial intelligence and statistics, pp. 276–283, 2003.
[63] Gardner, J., Pleiss, G., Weinberger, K. Q., Bindel, D., and Wilson, A. G., "Gpytorch: Blackbox matrix-matrix gaussian process inference with gpu acceleration", Adv Neural Inf Process Syst, Vol. 31, 2018.
[64] Al-Shoukairi, M., Schniter, P., and Rao, B. D., "A GAMP-based low complexity sparse Bayesian learning algorithm", IEEE Transactions on Signal Processing, Vol. 66, No. 2, pp. 294–308, 2017.
[65] Zhou, W., Zhang, H. -T., and Wang, J., "An efficient sparse Bayesian learning algorithm based on Gaussian-scale mixtures", IEEE Transactions on Neural Networks and Learning Systems, 2021.
[66] Roth, V., "The generalized LASSO", IEEE Trans Neural Netw, Vol. 15, No. 1, pp. 16–28, 2004.
[67] Titsias, M., "Variational learning of inducing variables in sparse Gaussian processes", in Artificial intelligence and statistics, pp. 567–574, 2009.
[68] Hensman, J., Matthews, A., and Ghahramani, Z., "Scalable variational Gaussian process classification", in Artificial Intelligence and Statistics, pp. 351–360, 2015.
[69] Hensman, J., Fusi, N., and Lawrence, N. D., "Gaussian processes for Big data", in Proceedings of the Twenty-Ninth Conference on Uncertainty in Artificial Intelligence, pp. 282–290, 2013.
[70] Wilson, A. G., Hu, Z., Salakhutdinov, R. R., and Xing, E. P., "Stochastic variational deep kernel learning", Advances in Neural Information Processing Systems, Vol. 29, 2016.
[71] Uhrenholt, A. K., Charvet, V., and Jensen, B. S., "Probabilistic selection of inducing points in sparse Gaussian processes", in Uncertainty in Artificial Intelligence, pp. 1035–1044, 2021.