Syed Zulqarnain Hassan.

GitHub ↗

Live model [todo]

Tweet emotion detection, shipped as a pipeline.

Type a tweet; the model returns happiness or sadness with a probability, and every stage that produced it, from the hashed source file to the held-out metrics, is tracked by DVC.

Demonstration system trained on public tweets labelled happiness or sadness by crowd workers in 2016. It scores short English text only and is not a mental-health tool.

Receipts

Measured on the held-out split.

Every number on this page is read from a file the pipeline wrote; nothing is typed by hand. Each tile links to that file on GitHub.

Test split, n = [todo]

Try it

Score a tweet.

Real held-out tweets, chosen by the shipped model's own scores, are preloaded. Pick one or type your own; the API normalises the text, scores it, and returns a probability.

Presets load from the API: [todo]

Short English text. Before scoring, the pipeline lowercases it, strips URLs, mentions, numbers, and punctuation, drops stop words, and lemmatises what is left.

[todo]

POST /predict
The pipeline

Six stages, one command.

Each stage declares its deps, params, and outs in dvc.yaml, so the graph is explicit and every artifact is hashed. The cards run left to right in the order DVC runs them.

  1. ingest

    Copy the repository's CSV (download it if it is missing), verify its sha256, keep two labels, drop duplicate texts, split.

    outs
    • data/raw/tweet_emotions.csv
    • data/raw/train.csv
    • data/raw/test.csv
    • data/raw/fetch_manifest.json
  2. preprocess

    Normalise text: case, URLs, mentions, numbers, punctuation, stop words, lemmas.

    outs
    • data/processed/train.csv
    • data/processed/test.csv
  3. features

    Fit the vectoriser on train only and transform both splits.

    outs
    • data/features/train.npz
    • data/features/test.npz
    • data/features/feature_manifest.json
    • models/vectorizer.joblib
  4. train

    Cross-validate the candidates, fit the winner, save the full pipeline.

    outs
    • models/model.joblib
    • models/version.json
  5. evaluate

    Score the held-out split once; write metrics, top terms, and figures.

    outs
    • reports/metrics.json
    • reports/top_terms.json
    • reports/figures/
  6. presets

    Pick real held-out tweets by the shipped model's own scores, for this page.

    outs
    • configs/presets.json
Reproduce
dvc repro

dvc repro reruns only the stages whose deps or params changed and skips the rest. dvc dag prints the graph; dvc metrics show prints the metrics file.

Top terms

Words that move the score.

The vocabulary terms with the largest weights in the shipped model, read from the file the evaluate stage wrote. Yellow leans happiness, blue leans sadness.

Loads from reports/top_terms.json through the API. [todo]

[todo]

Run it yourself

Same service, one image.

The image on GHCR contains the model, the vectoriser, the version and metric receipts, the presets, and this page. The curl example posts the first preset above, so the payload is a real held-out tweet.

Docker
docker run -p 8000:8000 ghcr.io/zulqarnain-10/tweet-emotion-pipeline:latest
Curl, first preset
curl -X POST http://127.0.0.1:8000/predict -H "Content-Type: application/json" -d '[todo]'