Briefings are dated notes on recent papers or developments worth understanding. An edition may cover one item or a small group.
Each item explains the question, the reported result, why it matters, and the limits that a reader needs to see. It is more than a link list, but it stops short of the full background, derivation, and claim-by-claim scrutiny expected of a Review.
The date identifies the edition and controls its place in the collection. It does not mean that every source appeared that day, and it does not promise a fixed publishing schedule. Reusable background belongs in a Concept; a paper that needs a deeper examination can become a separate Paper Review.
Example
P1. Attention Is All You Need
| Authors | Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin |
|---|---|
| Venue | Advances in Neural Information Processing Systems 30 (2017) · arXiv:1706.03762v7 |
| Tags | AI/ML |
More details
Background
- 2017년 당시 신경망 기계번역을 포함한 sequence transduction의 주류는 recurrent neural network(RNN), 특히 LSTM·GRU 기반 encoder-decoder였다. Attention은 이미 번역 성능을 크게 높인 핵심 기법이었지만, 대체로 recurrent network와 함께 사용되었다 [1] [2].
- RNN은 앞 위치의 hidden state를 다음 위치가 순차적으로 받아야 하므로 한 문장 안의 위치들을 동시에 계산하기 어렵다. Convolutional sequence model은 위치별 계산을 병렬화할 수 있지만, 멀리 떨어진 두 위치를 연결하려면 여러 층을 거쳐야 한다 [3]. 이 논문의 출발점은 recurrence와 convolution을 모두 제거하고 attention만으로 encoder-decoder를 구성할 수 있는가라는 질문이다.
Summary
저자들은 recurrent layer와 convolutional layer를 사용하지 않고, multi-head attention과 position-wise feed-forward network를 쌓은 encoder-decoder 구조인 Transformer를 제안한다. 순서 정보를 recurrence로 전달하지 않으므로 입력 embedding에 positional encoding을 더하고, decoder에는 미래 token을 보지 못하도록 masking을 적용한다 [1].
핵심 연산인 scaled dot-product attention은 다음과 같이 정의된다.
가 query와 key의 적합도를 계산하고, scaling이 큰 차원에서 softmax를 gradient가 극도로 작아지는 영역으로 밀어 넣는 현상을 완화한다. Multi-head attention은 서로 다른 학습된 부분공간에서 이 연산을 병렬로 수행해 여러 위치와 관계를 함께 표현한다 [1].
WMT 2014 영어-독일어 번역에서 큰 Transformer 모델은 28.4 BLEU를 기록해 ensemble을 포함한 당시 최고 결과를 2 BLEU 이상 넘어섰다. 영어-프랑스어에서는 41.0 BLEU로 당시 최고 단일 모델을 능가했고, 저자들의 추산으로 그 모델의 4분의 1 미만 training cost를 사용했다. 큰 모델의 학습에는 P100 GPU 8개로 3.5일이 걸렸다 [1].
Discussion
- 이 논문의 핵심은 attention 자체를 처음 발명한 데 있지 않다. 기존에는 RNN encoder-decoder를 보조하던 attention을 모델의 주된 계산 구조로 바꾸고 recurrence와 convolution을 제거했다는 데 있다. 그 결과 학습 중 각 sequence 위치의 표현을 병렬로 계산할 수 있고, self-attention 한 층에서 임의의 두 위치 사이의 최대 경로 길이가 상수가 된다 [1].
- 다만 “병렬화”를 생성 전체가 병렬화된다는 뜻으로 읽으면 안 된다. 학습 시 masked self-attention은 target 위치들을 함께 계산할 수 있지만, 논문의 decoder는 추론할 때 이전에 생성한 token을 입력으로 사용하는 autoregressive 구조다. 저자들도 generation을 덜 순차적으로 만드는 문제를 후속 과제로 남겼다 [1].
- Global self-attention의 위치 간 연결에는 비용이 따른다. 논문이 제시한 층별 복잡도는 self-attention이 , recurrence가 이다. 저자들은 번역에서 흔한 조건에서는 self-attention이 유리하다고 설명하지만, 매우 긴 입력을 위해서는 attention 범위를 제한하는 방법이 필요하다고 보았다 [1].
- 이 예시는 2017 NeurIPS proceedings판을 기준으로 한다. 이 판이 직접 입증한 범위는 두 WMT 2014 기계번역 과제이며, 영어-프랑스어 결과를 41.0 BLEU로 보고한다. 2023년의 arXiv v7은 영어 constituency parsing 실험을 추가하고 초록과 결과표의 영어-프랑스어 수치를 41.8 BLEU로 제시하므로, 평가 범위와 수치는 판본을 함께 밝혀야 한다. 이미지·오디오·비디오로의 확장은 두 판본 모두에서 실험 결과가 아니라 향후 연구 방향이다 [1].
References
- [1] Vaswani et al., “Attention Is All You Need”, Advances in Neural Information Processing Systems 30 (2017) · arXiv:1706.03762v7.
- [2] Bahdanau, Cho, and Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate”, International Conference on Learning Representations (2015).
- [3] Gehring et al., “Convolutional Sequence to Sequence Learning”, Proceedings of the 34th International Conference on Machine Learning (2017).
Research practice describes the shared source-checking and approval workflow.