In Part 1, we cleared the historical board. We examined how carbon-14 dating places MS 408 squarely in the early 15th century (c. 1404–1438), dismantled modern forgery theories, and situated its “bizarre” illustrations within late-medieval workshop culture and the Doctrine of Signatures.
Now comes the deeper, more frustrating half of the puzzle: the text itself.
For over a century, brilliant cryptographers, WWII codebreakers (including William Friedman), and modern computational linguists have attacked the manuscript with single-language hypotheses. They have claimed to find pure Latin, Hebrew, Proto-Romance, High German, or Old Turkish. Every single one of these attempts ultimately collapses under peer review. Why? Because the script behaves like natural language in one metric, but shatters how natural human languages work in almost every other.

The Linguistic Paradox: Why Standard Decipherment Fails
To understand why the Voynich Manuscript resists standard translation, you have to look at two contradictory statistical realities:
1. Zipf’s Law Holds True
If you take all ~38,000 words in the manuscript and plot their frequency distribution on a logarithmic graph, they form a near-perfect slope. The most frequent word appears roughly twice as often as the second, three times as often as the third, and so on. This mathematical fingerprint (Zipf’s Law) is the universal hallmark of authentic human communication. MS 408 is not the random, chaotic babbling of a madman.
In linguistics, Zipf’s Law means that a tiny handful of words are used constantly, while the vast majority of words are rarely used at all.
Specifically, the most common word in a language is used twice as often as the second most common word, three times as often as the third, and so on.
For example, in English:
🥇 Rank 1 (“the”) appears about 7% of the time.
🥈 Rank 2 (“of”) appears half as much, about 3.5% of the time.
🥉 Rank 3 (“and”) appears a third as much, about 2.3% of the time.
Essentially, human language is highly predictable: we rely heavily on a very small shortcut vocabulary to do most of our heavy lifting.
2. The Entropy Problem & Extreme Repetition
Despite obeying Zipf’s Law, the internal behavior of the words breaks natural language syntax:
- Unnaturally Low Conditional Entropy: Once you know the first two characters of a Voynich word, predicting the subsequent letters is mathematically trivial—far more predictable than in Latin, Italian, or German.
- The “Lego-Block” Combinatorial Trap: Words are not built from flexible vocabularies. They are assembled by snapping together a small kit of prefixes (qok-, ol-, ot-) with repeating root cores and terminal suffixes (-edy, -eey, -dain).
- Extreme Word Self-Similarity: Natural prose rarely repeats near-identical words sequentially. In Voynichese, patterns like qokedy qokedy or slight morphemic shifts appear continuously, alongside baffling “triplications” (such as sheol. sheol. sheol. on Folio 104v).
When scholars force a natural language dictionary onto this structure, it falls apart. If you assign a single Latin word to qokedy, that word ends up appearing five times on the same page next to a flower, a star, and a bathtub.
The Working Breakthrough: The Dual-Layer Hypothesis
During our collaborative computational modeling sessions with Gemini, we stepped away from searching for a single hidden prose language. Instead, we analyzed the text through the lens of a 15th-century Mediterranean/Alpine trade nexus.
What emerged is the Dual-Layer Model—a realization that the manuscript’s text is not uniform, but partitioned into two distinct structural components:
| THE DUAL-LAYER ARCHITECTURE |
| LAYER 1: The Procedural Engine (High-Frequency Shorthand) • Highly repetitive, prefix-driven, low entropy. • Functions as an operational command syntax (boil, steep, route). • Built from standardized 15th-century scribal tachygraphy. |
| LAYER 2: The Inset Lexicon (Low-Frequency Trade Loanwords) • Rare outliers and hapax legomena (< 5 occurrences). • Unencrypted proper nouns and materials from trade Lingua Franca. • Spoken loanwords written phonetically (Arabic, Romance, Germanic). |
Layer 1: The Procedural Engine
The bulk of the running text is not narrative prose; it is a highly compressed, prefix-driven operational shorthand designed for an apothecary laboratory and bathhouse facility.
Paleographically, these glyphs are stylized monograms adapted from 15th-century Latin scribal abbreviations (tachygraphy) and notary marks:
| Structural Element | Paleographic Origin | Operational Meaning | Real-World Application |
| ol- | Latin oleum / MHG oli- | Fluid / Distillate Base | Liquid medium, oils, infused bath waters, vessel capacity. |
| ot- | Slavic otъ- (outflow) / OHG ot- | Expressed Yield / Outflow | Active extraction, juice expression, pipe drainage lines. |
| qok- | Latin ligature con-coquere | Thermal / Batching Verb | Heating, boiling, steeping, regulating bath temperature. |
| -eey | Loop core + terminal tail | Conduit / Duration | Physical plant stalk, bath pipe, or sustained flow command. |
| sain-d | Latin sanus / Old French sain | Purified State | Clean, settled, uncorrupted base material. |
| Gallows (p-, f-) | Scriptorium section keys | Header / Unit Standards | Batch start, weight tallies, duration units (hours/passes). |
This explains the low entropy: an operator logging boiler cycles, steeping vats, and pipe valves uses the same formulaic commands over and over.

Layer 2: The Inset Lexicon (Wanderwörter)
If Layer 1 is the grammatical engine, Layer 2 provides the physical nouns.
When an apothecary or bathmaster needed to record a specific foreign plant, a specialized resin, or an architectural feature, they did not invent a code word. They recorded the spoken trade term phonetically—resulting in Wanderwörter (wandering words) that traveled trans-continental trade routes:
- daira (Folio 86v, Rosettes Map): Directly matches Arabic/Persian dā’ira (دائرة), meaning circular enclosure, ring, or circuit—labeling the circular causeway connecting the bath towers.
- ayn / ain (Folio 86v): The Semitic root ‘Ayn (عين), the universal historical term for an eye, natural water spring, or source head.
- sard-a (Folio 33r): Demonstrates liquid consonant metathesis, pointing to Sidra (سدرة), the Mediterranean Lote tree widely documented in Islamic and Southern European medicine for therapeutic, foaming herbal baths.
- otolora (Folios 88r and 84r): An agglutinative compound bridging Basque/Pyrenean ote (gorse/broom bush) and lore (flower). On Folio 88r, it names the harvested flower cutting; on Folio 84r, it reappears directly labeling the overhead canopy pipe dispensing that exact floral douse to bathers!

The Synthesis in Action
When you decouple the repeating procedural frame from the specific trade insets, the text stops looking like an impenetrable cipher.
Take a line from Folio 34v:
Raw Transcription: p-or-ain ot-ol-dy sardo qok-eey sain-a ol-sain
Decoded Operational Reading: “Primary extraction stock of the yield (p-or-ain). Fully process the liquid extract (ot-ol-dy) of the Sardonia / acrid herb (sardo); boil through the conduit channel (qok-eey) until reaching a purified state (sain-a) in a pure fluid medium (ol-sain).”
The syntax dictates how the operation runs (qok-, ot-ol-), while the unencrypted inset (sardo) tells the practitioner what is in the vat. In Part 3, we will put this Dual-Layer Engine directly into the workshop: stepping through the physical folios to watch raw botanical extractions on Folio 2r and Folio 88r flow straight into the thermal plumbing and multi-tier soaking pools of Folios 78r and 84