Constraints as Method: Notes on a Vector Database for Restricted Hardware
October 2025. A MacBook Pro M1, sixteen gigabytes of RAM. Qwen3.5 9B loaded through Ollama, Gemma running on the same hardware in alternate sessions. I was interested in retrieval augmented generation against a local corpus, executed without remote dependency. A vector database was required. I installed Qdrant as the default candidate.
The result was immediate: the language model occupied approximately five gigabytes of resident memory, Qdrant claimed three, and the operating system starved the remainder. Background processes terminated. The browser, opened to consult documentation, refused to allocate the requested pages. The configuration was not viable.
I had used Qdrant before. In previous contexts the question of footprint had not arisen. Memory was abundant. The October session forced a different reading of the same software. Vector retrieval, as it was being shipped in late 2025, assumed a hardware envelope I did not have. Cloud deployment, dedicated servers, abundant RAM. The personal laptop running a local model was outside the envelope.
I did not have a budget for the alternative. A new NVIDIA Spark configuration was approximately four thousand euros at the time. A Mac Studio in the relevant configuration was comparable. Neither was within reach. The machine I had was a 2021 MacBook Pro that worked well, and that I would prefer to keep using for several more years. The constraint was not ideological. It was the constraint of a person who likes to build things on the hardware they already own, and who does not see the case for replacing it.
The literature on vector search is consistent on the architectural question. The dominant indexing structure of the period is HNSW, the Hierarchical Navigable Small World graph proposed by Malkov and Yashunin in 2018 [1]. Its performance characteristics are well documented and its adoption near universal. The structure has one consequential property. The graph and the underlying vectors are held in resident memory in their entirety. For a corpus of one million vectors at typical embedding dimensions, this produces a memory footprint in the gigabyte range. The literature also documents alternatives. Vamana, introduced by Subramanya and colleagues in 2019 under the name DiskANN [2], is structurally similar to HNSW but designed for residence on secondary storage. The graph is paged. The vectors are accessed through disk reads. The resident set is bounded.
The alternatives were not, in late 2025, widely implemented in production systems intended for general use. A small number of research artefacts existed. A larger number of forks were available with incomplete characteristics. No system in the surveyed set combined the storage profile of Vamana with the operational properties expected from a modern database. Persistent transactions. Crash recovery. A wire protocol. Multiple language bindings. The gap was not a research gap. It was an engineering gap. The structures existed in papers and partial implementations. They had not been assembled into a system one could operate.
I decided to build one. The decision was not the product of market analysis, but of a personal preference for the work that takes place when constraints are imposed deliberately. I had spent the preceding months producing Gleam libraries and small applications under no particular external pressure. The vector database problem presented a constraint profile of unusual specificity. A laptop with sixteen gigabytes of memory. A coexistence requirement with a five gigabyte language model. A target footprint in the low hundreds of megabytes. A latency target in the low millisecond range. The combination defined a small region of the design space, and the small region had not been explored in production form.
The framing was not new. The history of computing contains repeated examples of useful work produced under self-imposed constraint. SQLite operates on the assumption that a database engine should run within a process boundary and require no administration. Redis operates on the assumption that the data set fits in memory and that durability is configurable rather than mandatory. Both systems make architectural choices that would be incorrect in a different envelope and that are correct in their declared envelope. The literature on constraint as a creative method is established. Stiny and Gips on shape grammars [3], Stravinsky on the productive function of imposed form [4], Brooks on the disciplinary effect of fixed-size memory budgets [5]. The principle is consistent. A budget produces structure.
I gave the system the name skeg. The skeg is the fixed component at the rear of a vessel’s keel. It neither propels nor steers; it determines that the line held by the vessel remains the line intended. The component is structural rather than performative. The name was selected on this basis. The system does not produce embeddings. It does not perform inference. It holds the retrieval substrate in a stable configuration so that the components above it operate as expected.
Work began in October 2025. The first commit established a workspace of empty crates and a set of design documents recording the hypotheses to be tested. The hypotheses included the choice of indexing structure, the choice of quantisation scheme, the choice of persistence strategy, the choice of wire protocol, the choice of concurrency model. Each hypothesis was scheduled for falsification. The method, recorded in the second post in this series, was to define a gate for each hypothesis. The gate specified the conditions under which the hypothesis would be considered confirmed. Implementation work was authorised only after the gate was passed. The expected outcome of any given gate, given the prior literature and my own estimation, was approximately even. Half the hypotheses would survive contact with measurement. Half would not.
The actual outcome was less favourable. Of the twelve hypotheses tested in the eight months that followed, eleven were falsified. The single confirmed hypothesis became the load-bearing component of the shipped system. The eleven failures produced the architecture, in the negative space. What the system is, is the residue of what it could not be.
The subsequent posts in this series record the gates, the falsifications, and the resulting structure. The intent is documentary. The structure of the work is more legible in retrospect than it was at any point during execution. I am recording the legibility here so that the next instance of the same exercise, by me or by someone else, may begin with a clearer view of which hypotheses are worth testing and which are not.
The system is available under Apache 2.0. The benchmark numbers, the architectural choices, and the residual limitations are recorded in the fourth post of this series. The intervening posts record the method.
References
[1] Malkov, Yury A.; Yashunin, Dmitry A. “Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs.” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 4, pp. 824–836, 2020. Preprint: arXiv
.09320, 2016.[2] Subramanya, Suhas Jayaram; Devvrit; Kadekodi, Rohan; Krishaswamy, Ravishankar; Simhadri, Harsha Vardhan. “DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node.” Advances in Neural Information Processing Systems 32, NeurIPS 2019.
[3] Stiny, George; Gips, James. “Shape Grammars and the Generative Specification of Painting and Sculpture.” Information Processing 71, North-Holland, 1972.
[4] Stravinsky, Igor. Poetics of Music in the Form of Six Lessons. Harvard University Press, 1947. The relevant passage on the productive function of constraint appears in the third lesson, “The Composition of Music.”
[5] Brooks, Frederick P. The Mythical Man-Month: Essays on Software Engineering. Addison-Wesley, 1975. The relevant material on memory budgets and architectural discipline appears in chapter 9, “Ten Pounds in a Five-Pound Sack.”