Text Classification with Machine Learning: Building Your First Classifier
Text classification — assigning predefined categories to pieces of text — is one of the most practical and widely deployed applications of natural language processing. Spam detection, sentiment analysis, topic categorization, language identification, intent detection in chatbots, and content moderation all rely on text classification at their core. Building a text classifier has become significantly more accessible in recent years through high-quality Python libraries and pre-trained models, but understanding the fundamentals helps you choose the right approach, collect the right training data, and evaluate your model correctly.
Feature Extraction: Converting Text to Numbers
Machine learning algorithms operate on numerical feature vectors, not raw text. Converting text to numbers is the feature extraction step, and the choice of representation significantly affects classifier performance. Bag of Words (BoW) represents a document as a vector of word counts, ignoring word order. TF-IDF (Term Frequency-Inverse Document Frequency) weights each word by how often it appears in the document relative to how common it is across all documents, giving higher weight to discriminative terms and lower weight to common words. N-gram features extend BoW to capture short word sequences: "not good" as a bigram carries different sentiment than the individual words "not" and "good." Modern approaches use pre-trained embeddings from transformer models as features, capturing semantic meaning that sparse BoW representations cannot represent. The choice depends on dataset size — transformer embeddings require substantial training data to fine-tune effectively, while TF-IDF classifiers can work well with only a few hundred labeled examples.
Classical Classifiers: Naive Bayes and SVMs
Naive Bayes is the traditional baseline classifier for text classification, and it performs remarkably well for its simplicity. It estimates the probability of each class given the observed word frequencies, using Bayes' theorem with the "naive" assumption that words are conditionally independent given the class. This assumption is technically wrong — word co-occurrences are clearly not independent — but it is wrong in a way that usually does not matter much in practice. Naive Bayes trains very quickly, handles high-dimensional sparse features naturally, and is highly interpretable. Support Vector Machines (SVMs) with TF-IDF features are a step up in accuracy for most tasks, finding the maximum-margin hyperplane separating classes in the high-dimensional feature space. SVMs are slower to train but typically outperform Naive Bayes, especially when training data is limited.
Training Data: Quality, Quantity, and Class Balance
The quality and quantity of labeled training data is usually the dominant factor in text classifier performance — more important than the choice of algorithm. For classical classifiers, a few hundred labeled examples per class can produce useful models; for fine-tuning transformer models, a few thousand examples per class is typically needed. Class balance matters: a classifier trained on 90% negative and 10% positive examples will learn to predict negative almost always, achieving 90% accuracy while being useless for detecting positive cases. Address imbalance through oversampling the minority class, undersampling the majority class, or adjusting class weights in the loss function. For annotation projects, inter-annotator agreement — the degree to which different human annotators label the same text the same way — is a critical quality metric. Low agreement indicates ambiguous label definitions that will confuse any classifier trained on the data.
Evaluation Metrics: Beyond Accuracy
Accuracy (fraction of correct predictions) is intuitive but can be misleading when class frequencies are unequal. For imbalanced datasets, metrics that treat classes symmetrically are more informative. Precision measures what fraction of items labeled as positive are truly positive; recall measures what fraction of truly positive items were labeled as positive. The F1 score is the harmonic mean of precision and recall, providing a single balanced metric. For multi-class problems, these metrics can be averaged across classes (macro average) or weighted by class frequency (weighted average). The confusion matrix visualizes where the classifier makes errors — which classes are being confused with which — and is essential for identifying systematic error patterns that aggregate metrics can hide.
Want to go further? Continue with the other LibriText explainers, or suggest a topic you would like covered.