NanoRush Chat
NanoRush Chat is a 283M parameter GPT-style causal language model fine-tuned for conversational AI.
Github- https://github.com/Amogh1221/NanoRush
Live- https://nano-chat-web.vercel.app
Model Details & Configuration
| Detail | Value |
|---|---|
| Parameters | 283M |
| Architecture | GPT-2 style |
| Precision | FP16 / BFloat16 |
| Context Window | Up to 4096 tokens |
| Vocabulary Size | 32,768 |
| Embedding Dimension (n_embd) | 768 |
| Number of Heads (n_head) | 12 |
| Number of Layers (n_layer) | 36 |
| Base Model | Custom pre-trained |
| Pre-Training Dataset | HuggingFaceTB/cosmopedia |
| Fine-tuning Dataset | HuggingFaceTB/smoltalk |
Evaluation Results
The model was evaluated using standard zero-shot accuracy metrics.
| Groups / Tasks | Version | n-shot | Metric | Value | Stderr |
|---|---|---|---|---|---|
| mmlu | 2 | 0 | acc | 0.2297 | ± 0.0035 |
| - humanities | 2 | 0 | acc | 0.2438 | ± 0.0063 |
| - other | 2 | 0 | acc | 0.2375 | ± 0.0076 |
| - social sciences | 2 | 0 | acc | 0.2184 | ± 0.0074 |
| - stem | 2 | 0 | acc | 0.2119 | ± 0.0073 |
| arc_challenge | 1 | 0 | acc | 0.2295 | ± 0.0123 |
| hellaswag | 1 | 0 | acc | 0.3116 | ± 0.0046 |
| truthfulqa_mc2 | 3 | 0 | acc | 0.4320 | ± 0.0153 |
| winogrande | 1 | 0 | acc | 0.5107 | ± 0.0140 |
Usage
This model has been exported to be fully compatible with the Hugging Face transformers library.
You can load it using the standard AutoModelForCausalLM pipeline.
Installation
Make sure you have the latest version of the transformers and torch libraries installed:
pip install torch transformers
Example Code
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from transformers import StoppingCriteria, StoppingCriteriaList
from transformers.generation.streamers import TextIteratorStreamer
import threading
# Load the model and tokenizer from Hugging Face
model_id = "Amogh1221/nano-chat"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float32,
device_map="cpu"
)
system_prompt =
"""
You are NanoRush, an AI assistant, you are a helpful, respectful, and intelligent conversational
partner. You must never pretend to be a human, and you must carefully pay attention to the conversation history.
"""
class StopOnUser(StoppingCriteria):
def __init__(self, prompt_length):
self.prompt_length = prompt_length
def __call__(self, input_ids, scores, **kwargs):
generated_tokens = input_ids[0][self.prompt_length:]
tail = tokenizer.decode(generated_tokens[-10:])
return "\nUser:" in tail or "User:" in tail
# Format your prompt
prompt = f"System: {system_prompt}\\n\\nUser: What is Quantum Computing?\\nAssistant:"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
stop_criteria = StoppingCriteriaList([StopOnUser(prompt_length=inputs["input_ids"].shape[1])])
streamer = TextIteratorStreamer(tokenizer, skip_prompt=True, skip_special_tokens=True)
generation_kwargs = dict(
**inputs,
max_new_tokens=512,
temperature=0.7,
top_k=50,
top_p=0.9,
do_sample=True,
repetition_penalty=1.15,
pad_token_id=tokenizer.eos_token_id,
prompt_lookup_num_tokens=3,
stopping_criteria=stop_criteria,
streamer=streamer,
)
# Run generation in a background thread
thread = threading.Thread(target=model.generate, kwargs=generation_kwargs)
thread.start()
print("Assistant: ", end="")
for text in streamer:
print(text, end="", flush=True)
print()
- Downloads last month
- 65
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support