рдореБрдЦреНрдп рдордЬрдХреБрд░рд╛рдХрдбреЗ рдЬрд╛
JobCannon
рд╕рд░реНрд╡ рдХреМрд╢рд▓реНрдпреЗ

Natural Language Processing (NLP)

Teach computers to understand text: sentiment, translation, LLMs

тмв рд╢реНрд░реЗрдгреА 2рддрд╛рдВрддреНрд░рд┐рдХ
+$40k-
рдкрдЧрд╛рд░рд╛рд╡рд░реАрд▓ рдкрд░рд┐рдгрд╛рдо
12 рдорд╣рд┐рдиреЗ
рд╢рд┐рдХрдгреНрдпрд╛рд╕ рд▓рд╛рдЧрдгрд╛рд░рд╛ рд╡реЗрд│
рдХрдареАрдг
рдХрд╛рдард┐рдгреНрдп
11
рдХрд░рд┐рдЕрд░реНрд╕
рдПрдХрд╛ рджреГрд╖реНрдЯрд┐рдХреНрд╖реЗрдкрд╛рдд

NLP is teaching computers to understand human language through preprocessing, embeddings, and transformer models (BERT, GPT). Career path: NLP Engineer L1 (sentiment analysis, text classification, $100-150k) тЖТ L2 (transformers, fine-tuning, $160-220k) тЖТ L3/Research Scientist (LLM pretraining, RLHF, $200-350k). Transformers replaced RNNs as the dominant architecture in 2018-2020; GPT/BERT skills command premium salaries due to LLM boom. Tech stack: Python + spaCy/NLTK for preprocessing, Hugging Face Transformers library, PyTorch for training, LangChain for production LLM apps, vector databases (Pinecone/Weaviate) for semantic search and RAG.

Natural Language Processing (NLP) рдореНрд╣рдгрдЬреЗ рдХрд╛рдп

NLP = AI that understands human language. Sentiment analysis, translation, chatbots, LLMs (GPT, BERT). High-demand ML specialty (ChatGPT boom). L1: Text preprocessing, sentiment analysis, word embeddings

ЁЯФз рд╕рд╛рдзрдиреЗ рдЖрдгрд┐ рдкрд░рд┐рд╕рдВрд╕реНрдерд╛
Hugging FacespaCyNLTKTransformersPyTorchBERTGPTLangChainOpenAI APIPineconesentence-transformersLLaMAWeaviate

ЁЯУЛ рд╕реБрд░реВ рдХрд░рдгреНрдпрд╛рдкреВрд░реНрд╡реА

ЁЯТ░ рдкреНрд░рджреЗрд╢рд╛рдиреБрд╕рд╛рд░ рдкрдЧрд╛рд░

рдкреНрд░рджреЗрд╢рдЬреНрдпреБрдирд┐рдпрд░рдордзреНрдпрдорд╕реАрдирд┐рдпрд░
USA$120k$180k$280k
UK┬г70k┬г110k┬г160k
EUтВм75kтВм120kтВм180k
CANADAC$125kC$185kC$290k

тЪЦ рдпрд╛рдВрдЪреНрдпрд╛рд╢реА рддреБрд▓рдирд╛ рдХрд░рд╛

тЭУ FAQ

Transformers vs RNNs/LSTMs, why did transformers win?
RNNs (LSTM, GRU) process sequences one token at a time, can't parallelize, forget long-range context. Transformers (BERT, GPT) process all tokens in parallel via self-attention, capture long-range dependencies, train 10-100x faster. Attention is All You Need (2017) proved it. By 2020: transformers = industry standard. Learning RNNs = learning history; building production NLP = transformers only.
BERT vs GPT, which should I learn first?
BERT = bidirectional, pretrained on masked language modeling, best for classification/understanding tasks (sentiment, entity extraction, Q&A). GPT = unidirectional (left-to-right), pretrained on causal language modeling, best for generation (chatbots, summarization, translation). Learn BERT first (conceptually simpler, https://huggingface.co/course/chapter1), then GPT. In 2026: fine-tune BERT for custom classifiers, use GPT for chat/generation via API.
Fine-tuning vs RAG (Retrieval Augmented Generation) vs prompt engineering, when to use each?
Prompt engineering = free, fast, no training (try first). RAG = retrieve relevant documents from a vector database, feed to LLM context, good for knowledge-intensive tasks (Q&A over company docs), no retraining. Fine-tuning = expensive (GPUs), slow (hours-days), but fits the model to your data/style. Pick: (1) try prompt engineering, (2) if context window insufficient, add RAG, (3) if LLM still fails, fine-tune. 80% of use cases = prompt engineering + RAG.
How do I host/deploy an LLM myself vs using an API?
API (OpenAI, Anthropic) = $0.01-0.1 per 1k tokens, lowest latency, no infra. Self-host small LLMs (Llama 7B, Mistral) = $0.50-5/hour GPU, latency 100-500ms, full control, privacy. Rule: API for prototyping and scaling (chat, content generation), self-host for privacy-critical apps (healthcare, finance) or if volume > 10M tokens/month. 2026 trend: smaller specialized models (MistralAI) on your infrastructure, not giant models via API.
Vector databases and embeddings, Pinecone vs Weaviate vs building DIY?
Embeddings = convert text to numbers (768-1536 dims), enable semantic search. Pinecone = managed (easiest, $0.10-1/month), Weaviate = self-host (free, complex), DIY = index with NumPy (only for <10k docs). For production: Pinecone if budget available, Weaviate if on-premise required, DIY only for prototypes. All use sentence-transformers (all-MiniLM-L6-v2) for encoding and cosine similarity for retrieval.
Multilingual NLP, how hard is it to support multiple languages?
Monolingual models (English BERT) fail on other languages. Solutions: (1) multilingual BERT (mBERT, XLM-RoBERTa) for 100+ langs, lower quality, (2) language-specific models (French BERT, German BERT) for top langs, better quality, (3) translate to English (lossy but works). Recommendation: mBERT for MVP, switch to language-specific for each supported language in production. Don't try to build language-universal; use existing multilingual checkpoints from Hugging Face.
How do I evaluate NLP models, metrics beyond accuracy?
Classification: precision/recall/F1 (imbalanced), ROC-AUC (ranking). Generation (summarization, translation): ROUGE (n-gram overlap), BLEU (precision on n-grams), human evaluation (expensive). Token classification (NER): micro/macro F1 (per-token). Semantic similarity: cosine similarity, human correlation. NEVER use accuracy alone for text tasks, almost all are imbalanced. For LLMs: use LLM-as-judge (ask GPT to score quality) for generation tasks.

рд╣реЗ рдХреМрд╢рд▓реНрдп рддреБрдордЪреНрдпрд╛рд╕рд╛рдареА рдпреЛрдЧреНрдп рдЖрд╣реЗ рдХрд╛, рдпрд╛рдЪреА рдЦрд╛рддреНрд░реА рдирд╛рд╣реА?

рдХрд░рд┐рдЕрд░ рдореЕрдЪ рдХрд░реВрди рдкрд╛рд╣рд╛ тАФ рдЖрдореНрд╣реА рдпреЛрдЧреНрдп рдорд╛рд░реНрдЧ рд╕реБрдЪрд╡реВ.

рдорд╛рдЭреНрдпрд╛рд╕рд╛рдареА рд╕рд░реНрд╡реЛрддреНрддрдо рдХреМрд╢рд▓реНрдпреЗ рд╢реЛрдзрд╛ тЖТ

рддреБрдордЪрд╛ рдЖрджрд░реНрд╢ рдХрд░рд┐рдЕрд░ рдорд╛рд░реНрдЧ рд╢реЛрдзрд╛

реи,релреирез рдХрд░рд┐рдЕрд░рдордзреНрдпреЗ рдХреМрд╢рд▓реНрдпрд╛рдВрд╡рд░ рдЖрдзрд╛рд░рд┐рдд рдЬреБрд│рдгреА. рдореЛрдлрдд, ~3 рдорд┐рдирд┐рдЯреЗ.

рдХрд░рд┐рдЕрд░ рдореЕрдЪ рдХрд░реВрди рдкрд╛рд╣рд╛ тАФ рдореЛрдлрдд тЖТ