Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #426

Seq2Seq Model

The Seq2Seq (Sequence-to-Sequence) model is the complete, assembled system — encoder, context vector, and decoder working together — formalized as one end-to-end trainable architecture.

The Full Pipeline

\[ \mathbf{c} = \text{Encoder}(\mathbf{x}_1, \ldots, \mathbf{x}_T) \qquad \hat{\mathbf{y}}_1, \ldots, \hat{\mathbf{y}}_{T'} = \text{Decoder}(\mathbf{c}) \]

\(T\) and \(T'\) can be completely different — this is exactly the length-independence established in Encoder-Decoder Architecture. The decoder generates its output one token at a time, feeding each generated token back in as input for producing the next one, continuing until it produces a special "end of sequence" token.

Diagram — Autoregressive Decoding

context c <start> "J'aime" "les" "chiens" each generated word feeds back in as the next step's input

Generation is autoregressive — each output token depends on every token generated before it, plus the original context vector.

Code — A Simplified Full Decoding Loop

import torch

def generate(encoder, decoder, input_seq, start_token, end_token, max_len=20):
    h, c = encoder(input_seq)
    token = torch.tensor([[start_token]])
    output_tokens = []

    for _ in range(max_len):
        logits, h, c = decoder(token, h, c)
        next_token = logits.argmax(dim=-1)   # greedy decoding: pick the most likely token
        if next_token.item() == end_token:
            break
        output_tokens.append(next_token.item())
        token = next_token   # feed the generated token back in as the NEXT step's input

    return output_tokens

Training vs Inference — A Key Difference

During inference (as shown above), each step's input is the model's own previous prediction. During training, a different strategy — teacher forcing — is typically used instead, feeding the true previous token rather than the model's own (possibly wrong, especially early in training) prediction. This distinction, and the subtlety it introduces, is the entire subject of the next note.

Common Mistakes

  • Confusing the training-time and inference-time decoding procedures — they typically differ (teacher forcing during training, autoregressive generation during inference), and forgetting this distinction leads to a common category of subtle bugs.
  • Forgetting to handle the end-of-sequence token correctly — without a stopping condition, greedy decoding could run indefinitely (hence the max_len safeguard above).

Interview Relevance

Q: "Why is decoding in a Seq2Seq model called 'autoregressive'?" Each generated token becomes part of the input for generating the next token — the model's own previous outputs feed back into itself, one step at a time, rather than the entire output sequence being produced all at once. This mirrors the "auto-regressive" naming from statistics, where a value depends on its own previous values.

Practice Question

What would happen during inference if the decoder never learned to correctly predict an end-of-sequence token for a given input?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →