Burton Lancaster PRO
AI & ML interests
Recent Activity
Organizations
It indexes folders on your machine and answers plain-language or code questions with file paths, line ranges and the passages themselves.
RiverRider/motherlode-code-small-en-v0.1, a 33M-parameter code-search encoder, embeds them on your CPU. The first run makes six requests, all to Hugging Face, for the model; after that it makes none.
Measured once on an M2 Ultra through the MCP SDK's own client: the Sunstone repository, 57 files and 1,367 passages, woven in 20.2 s. Weaving it again took 0.014 s, and a search took 6 ms.
Claude Code:
claude mcp add weave -- npx -y blackwindow-weave-mcp --folder /path/to/repo
Cursor, VS Code, Claude Desktop and Windsurf run the same npx command from their MCP config.
npm: https://www.npmjs.com/package/blackwindow-weave-mcp
Registry: io.github.space-bacon/blackwindow-weave-mcp
Source: https://github.com/space-bacon/blackwindow-weave-mcp
Model: RiverRider/motherlode-code-small-en-v0.1
Licence: BUSL-1.1 with an internal-use grant, becoming Apache-2.0 on
15 September 2030.
Every figure you give reproduces from those files. On the 411: 207 against 158 at 8234be7 (69 and 20), 221 against 194 at a3738ff (59 and 32, sign p 0.0061), and 230 against 192 on the published path. Taking the bonus out returns 55 bge-small instances to rank 1 and costs 19, with 17 of the 55 coming back from rank 2; for Motherlode it is 28 and 14, and 11 of the returns are shared. The median raw spread of the top 10 hits at 8234be7 is 0.072 against 0.13. So 22 of the post's 49 was the bonus.
The post went up at 01:36 UTC on the 29th, while the page and Sunstone 0.1.5 still ran 8234be7's tool path. a3738ff, committed at 02:32, took the bonus out; Sunstone 0.1.6 went live with it at 02:46 and the page moved the same night. So "as it ships" held for about an hour. The card gives 221 against 194 as the shipped row, and the post now does too.
Both halves of your question, run on the 411 before writing this:
- After centring. The decomposition already holds it: without the raw rerank, the bonus goes on the centred dot the published path sorts on. It costs more there, not less: Motherlode 227 to 207 (12 better, 32 worse, sign p 0.0037) and bge-small 191 to 153 (17 and 55, sign p 0.00001).
- Scaled by spread. Registered before the run, with the rule that it comes back only if it helps both readers at sign p below 0.05: each query's bonus multiplied by that query's gap between its 1st and 10th raw cosines, over 0.14, so the full bonus lifts a passage at most from 10th to 1st for either reader. Motherlode goes 221 to 215 (5 better, 11 worse, sign p 0.21) and bge-small 194 to 190 (6 and 10, sign p 0.45). Motherlode leads 215 to 190.
So your mechanism holds: in each query's own units the bonus costs the two readers about the same, 6 and 4 files, where the fixed bonus cost them 14 and 36. Scaled or centred, it helps neither reader, so it stays off. On the weave's own questions it moved none of 54 in or out of the top 8.
Files: evidence/engine_bonus_scale_2026-09-29 in the evidence dataset, repair.json for your figures and summary.json for the scaled run, written by gates/engine_bonus_scale.py, whose docstring carries the rule.
RiverRider/motherlode-code-small-en-v0.1 is a 33.4M-parameter embedding model for code and text with bge-small-en-v1.5's architecture, so it drops in where bge-small runs. In a replay of Black Window's engine as it ships, on the 411 SWE-bench Verified issues its contamination check leaves, it ranks the file the fix changed first for 207, where bge-small does for 158: 69 issues better, 20 worse, exact sign test p below 0.00001.
One engine, three places:
- blackwindow.xyz, in a browser tab, with no account and no install
- Sunstone 0.1.5 on the Visual Studio Marketplace, which indexes the folders you open in VS Code for any model in the picker
- a GPU box you rent from the page on your own vast ai account, whose pairing line adds the same box to Sunstone
Model: RiverRider/motherlode-code-small-en-v0.1
Per-instance evidence: RiverRider/motherlode-code-small-en-v0.1-evidence
Page: https://blackwindow.xyz
Sunstone: https://marketplace.visualstudio.com/items?itemName=sunstonenorth.sunstone
Update, 2026-09-29: about an hour after this went up, engine a3738ff took the lexical bonus off the search path, and blackwindow.xyz and Sunstone 0.1.6 ship it. As it ships now, on the same 411: 221 against 194, 59 better and 32 worse, sign p 0.0061. The 207 against 158 above is the engine with the bonus, which cost bge-small more than Motherlode (see the comments).
Demo (+ source code): webml-community/DINOv3-video-tracking
This will revolutionize AI-powered video editors... which can now run 100% locally in your browser, no server inference required (costs $0)! ๐
How does it work? ๐ค
1๏ธโฃ Generate and cache image features for each frame
2๏ธโฃ Create a list of embeddings for selected patch(es)
3๏ธโฃ Compute cosine similarity between each patch and the selected patch(es)
4๏ธโฃ Highlight those whose score is above some threshold
... et voilร ! ๐ฅณ
You can also make selections across frames to improve temporal consistency! This is super useful if the object changes its appearance slightly throughout the video.
Excited to see what the community builds with it!
Did you ask your coding agent where something is handled, only to watch it grep, guess, and answer from the wrong file?
Did you edit a function, only to have the agent quote back the version from before your edit?
Sunstone is a free VS Code extension built for all three. It indexes the folders you have open and gives whichever model you already use, Copilot's Claude and GPT included, a real search over them. Before every search it re-reads whatever changed: a save, an unsaved edit, a file written from the terminal, a branch switch. #remember keeps a decision across chats. It can also put your own llama.cpp or vLLM server in the model picker. Everything stays on your machine, with no account and no key.
We measured it on all 500 issues in SWE-bench Verified. It named the right file first 229 times. A text search over the same files managed 58, and questions shuffled onto the wrong issues got 5. Every ranked list is published.
https://marketplace.visualstudio.com/items?itemName=SunstoneNorth.sunstone
Sunstone indexes the folders you have open and gives whichever model you are already using, Copilot's Claude and GPT included, a proper way to search them. It also lets you put your own servers in VS Code's model picker, if you want to.
Everything is indexed and held on your machine. There is no account and no key to hand over.
https://marketplace.visualstudio.com/items?itemName=SunstoneNorth.sunstone
Someone rolls an ankle on wet rock. Or there is a tick. Or the stream is silty and all you have is a filter and a cup.
If you loaded the tab before you left, the answers are still in your pocket.
https://blackwindow.xyz
Open it at home. Pick a model. Hit Load. Add it to the home screen.
Before you lose signal, drop in what you would actually need: a first-aid PDF, plant notes for the region, a map, a GPX, a meds list. The Weave indexes them on the device. Later questions pull the nearest passages back as notes.
After Load, the network can drop. The weights run in that tab, on that phone.
You can still drop new things in with no service. A photo of a plant, a rash, a trail junction. A voice note of symptoms. Audio is transcribed. An image is described. Both land in the Weave and you can ask about them there, offline.
It is a small model. Treat it as a checklist, not a doctor. It will not replace a beacon or a course. It will still answer when the map has no service.
Same trick on a plane, in a blackout, or with a file that should not leave the device.
Not a server with a policy. Your hardware, a window, a Load button.
Where would you actually use this?
https://blackwindow.xyz
Open the site, pick a model (about 0.6B to 8B), hit Load. The weights run in that tab, on that computer. After they load, the network can drop. The context window is a working set, auto-sized to that device, up to ~32K tokens.
Behind the window is the Weave. Every file, picture, recording, link, lookup, and reply is embedded as it arrives. Drop in audio and it is transcribed. Drop in an image and it is described. A question pulls the nearest passages back as notes. A long document is walked once so later questions can use the whole file, not the first pages.
Nothing leaves that tab unless you turn on live lookup or connect a rented GPU box, and the chat says so each time. Prompts can go to the box. Files and the Weave stay in the tab.
Console on that page: bw.ask, bw.search, bw.digest, bw.notes. A local relay exposes /v1/chat/completions on localhost so other tools on the same computer can talk to the tab. The tab polls the relay. That is the boundary.
Not a server with a policy. Your hardware, a window, a Load button.
If on mobile add to home-screen for best performance. If you break it lmk. It can serve a few hundred of you at a time before I have to buy a real server.
Building it turned up three things.
We graded 330 models on Korean and two axes collapsed.
Honorifics โ only 8.5% earn an A
Knowledge of Korean institutions โ 9.4%
Every other axis sits above 31%
Fluency hides it. A model can write clean, natural Korean and still attach an honorific to a coffee cup. Fluent and wrong at the same time is worse than obviously broken, because nobody catches it in review.
A 2023 model beats the 2026 flagships. gpt-3.5-turbo-16k scores a perfect 3.00. Korean cannot be inferred from release date, parameter count or English benchmarks โ it has to be measured, per model.
Quality, value and speed are three different models. Across five axes, the same model almost never takes two columns.
425 models, latency measured on 329 on a paid API, Korean graded on 330. Three languages, three currencies, daily refresh, open API, no key.
๐ https://huggingface.co/blog/ginigen-ai/openrouter-leaderboard ๐ ginigen-ai/open-router-leaderboard