Instructions to use Badhon/airline-intent-fasttext with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use Badhon/airline-intent-fasttext with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("Badhon/airline-intent-fasttext", "model.bin")) - Notebooks
- Google Colab
- Kaggle
airline-intent-fasttext
A 19-class intent classifier for airline customer-support messages written in code-mixed Bengali / Banglish (romanised Bengali) / English.
This is the recommended model for this domain. It outperforms the DistilBERT student trained on the same data by +17 points on the held-out human-written test set, while being 20x smaller and ~230x faster.
| artifact | size | use |
|---|---|---|
intent.ftz |
6.0 MB | ship this — quantized, negligible accuracy cost |
intent.bin |
81.5 MB | full precision, ~1.7 pts better |
Why fastText beats a transformer here
The corpus is transliterated Banglish, where the same word arrives spelled many
ways — koto / kotoo / kt. fastText's character n-grams (minn=3, maxn=5)
share weight across those spellings by construction. A subword-tokenised
transformer has to learn the equivalences from data it does not have enough of.
Intents (19)
greeting, goodbye, thanks, flight_status, booking_new, booking_manage,
checkin_boarding, baggage_policy, baggage_issue, refund_compensation,
fare_payment, airport_info, special_assistance, loyalty_program,
travel_documents, disruption, complaint, agent_request, out_of_scope
out_of_scope is the reject class — adjacent-service questions (train tickets,
hotels, visa agencies) that the airline does not handle.
Usage
import fasttext
from huggingface_hub import hf_hub_download
path = hf_hub_download("Badhon/airline-intent-fasttext", "intent.ftz")
model = fasttext.load_model(path)
labels, probs = model.predict("amar flight ta koto tay chare janaben?", k=3)
# ('__label__flight_status', ...) 0.94
intent = labels[0].removeprefix("__label__")
confidence = probs[0]
Inputs should be lowercased and stripped of trailing newlines; fastText treats a newline as a document boundary.
Confidence gating. Route to a human below ~0.6. Two known failure shapes sit above a 0.5 threshold and will pass it silently:
- adjacent-service questions leak into
booking_newwith high confidence (apnara ki train er ticket o katen→booking_new0.88, goldout_of_scope) - bare social tokens are low-confidence (
assalamu alaikum→greeting0.56)
out_of_scope recall on the synthetic split is 0.358. Do not rely on the reject
class alone to catch off-domain traffic.
Evaluation
Two evaluation sets. Rank on h-test — the hand-written adversarial holdout,
unseen during any tuning. The synthetic test split is generated from the same
templates as training and overstates every model.
| model | test | h-dev | h-test | p50 | size |
|---|---|---|---|---|---|
| fasttext (quantized) | 0.772 | 0.869 | 0.729 | 0.03 ms | 6.0 MB |
| fasttext (full) | 0.774 | — | 0.746 | 0.02 ms | 81.5 MB |
| distilbert-int8 (student) | 0.846 | 0.656 | 0.559 | 6.95 ms | 121.6 MB |
Full 120-item holdout accuracy 0.800, macro-F1 0.798.
Note DistilBERT scores higher on the synthetic split (0.846) and far lower on real phrasing (0.559). That divergence is the whole argument for ranking on h-test.
Intended use
Front-line intent routing for an airline support bot or IVR, with a confidence threshold and a human fallback. Not a decision system — it selects a reply template or a queue, nothing more.
- Downloads last month
- 13