NanoRush Chat

NanoRush Chat is a 283M parameter GPT-style causal language model fine-tuned for conversational AI.

Github- https://github.com/Amogh1221/NanoRush

Live- https://nano-chat-web.vercel.app

Model Details & Configuration

Detail Value
Parameters 283M
Architecture GPT-2 style
Precision FP16 / BFloat16
Context Window Up to 4096 tokens
Vocabulary Size 32,768
Embedding Dimension (n_embd) 768
Number of Heads (n_head) 12
Number of Layers (n_layer) 36
Base Model Custom pre-trained
Pre-Training Dataset HuggingFaceTB/cosmopedia
Fine-tuning Dataset HuggingFaceTB/smoltalk

Evaluation Results

The model was evaluated using standard zero-shot accuracy metrics.

Groups / Tasks Version n-shot Metric Value Stderr
mmlu 2 0 acc 0.2297 ± 0.0035
- humanities 2 0 acc 0.2438 ± 0.0063
- other 2 0 acc 0.2375 ± 0.0076
- social sciences 2 0 acc 0.2184 ± 0.0074
- stem 2 0 acc 0.2119 ± 0.0073
arc_challenge 1 0 acc 0.2295 ± 0.0123
hellaswag 1 0 acc 0.3116 ± 0.0046
truthfulqa_mc2 3 0 acc 0.4320 ± 0.0153
winogrande 1 0 acc 0.5107 ± 0.0140

Usage

This model has been exported to be fully compatible with the Hugging Face transformers library. You can load it using the standard AutoModelForCausalLM pipeline.

Installation

Make sure you have the latest version of the transformers and torch libraries installed:

pip install torch transformers

Example Code

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from transformers import StoppingCriteria, StoppingCriteriaList
from transformers.generation.streamers import TextIteratorStreamer
import threading

# Load the model and tokenizer from Hugging Face
model_id = "Amogh1221/nano-chat"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, 
    torch_dtype=torch.float32, 
    device_map="cpu"
)
system_prompt =
"""
You are NanoRush, an AI assistant, you are a helpful, respectful, and intelligent conversational
partner. You must never pretend to be a human, and you must carefully pay attention to the conversation history.
"""

class StopOnUser(StoppingCriteria):
    def __init__(self, prompt_length):
        self.prompt_length = prompt_length
    def __call__(self, input_ids, scores, **kwargs):
        generated_tokens = input_ids[0][self.prompt_length:]
        tail = tokenizer.decode(generated_tokens[-10:])
        return "\nUser:" in tail or "User:" in tail

# Format your prompt
prompt = f"System: {system_prompt}\\n\\nUser: What is Quantum Computing?\\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

stop_criteria = StoppingCriteriaList([StopOnUser(prompt_length=inputs["input_ids"].shape[1])])
streamer = TextIteratorStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)

generation_kwargs = dict(
    **inputs,
    max_new_tokens=512,
    temperature=0.7,
    top_k=50,
    top_p=0.9,
    do_sample=True,
    repetition_penalty=1.15,
    pad_token_id=tokenizer.eos_token_id,
    prompt_lookup_num_tokens=3,
    stopping_criteria=stop_criteria,
    streamer=streamer,
)

# Run generation in a background thread
thread = threading.Thread(target=model.generate, kwargs=generation_kwargs)
thread.start()

print("Assistant: ", end="")
for text in streamer:
    print(text, end="", flush=True)
print()
Downloads last month
65
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Amogh1221/NanoRush 1