Nutristatic Usage Guide

What It Is

Nutristatic is a serverless rewrite of Nutrimatic, Dan Egnor's pattern-matching word-search tool for puzzle solving and construction. Like Nutrimatic, it matches regular-expression-style patterns (with extensions such as anagrams and intersections) against a dictionary of every word and phrase that occurs in Wikipedia, ordering results by how common they are — so "reasonable" matches come first and "stretchy" ones later.

Word-pattern tools like OneLook and anagram tools like the Internet Anagram Server answer questions like "what six-letter word starts with 'kr' and ends with 'w'?" Nutrimatic's twist — inherited here — is that the dictionary isn't a fixed word list but the phrases of Wikipedia itself, so celebrity names, place names, catchphrases and idioms all match, results are ranked by corpus frequency, and the engine can splice words together to find phrases nobody put in a dictionary (famously "subject of blood and whiskey"), without being told where the word breaks are.

How Nutristatic Differs from Nutrimatic

Pattern Language Syntax

The query language is unchanged from Nutrimatic: it is based on regular expressions, in particular POSIX extended regular expressions (but without POSIX character classes).

The text (words and phrases) being matched has been normalized so that it is all lowercase and contains no punctuation -- only letters, numbers, and spaces. Most punctuation is turned into spaces, but apostrophes are simply removed, so "Fleur-de-lis" shows up as "fleur de lis", and "I'm Jack's total lack of surprise." is "im jacks total lack of surprise". Spaces are matched automatically by default; see the section on "quoted phrases" below for details.

Because only lowercase letters, digits and spaces show up in the text, uppercase letters and punctuation symbols are available to mean other things in queries:

Character Classes

For convenience, some uppercase letters and punctuation marks refer to commonly used classes of characters:

A - any alphabetic character, equivalent to [a-z]
C - any consonant (including y), equivalent to [bcdfghjklmnpqrstvwxyz]
V - any vowel (excluding y), equivalent to [aeiou]
_ (underscore) - any letter or number, equivalent to [a-z0-9]
- (hyphen) - an optional space, equivalent to ( ?)
# (number sign) - any digit, equivalent to [0-9]
[^...] - negated character class, matching any character in the alphabet not listed (e.g. [^aeiou])

Note that the hyphen and the underscore (as opposed to the standard regular expression dot ".") are only really useful inside "quoted phrases", though technically they are allowed anywhere.

Repeat quantifiers ({n}, {min,max}) are supported with limits up to 255.

"Quoted Phrases"

By default, spaces may be inserted anywhere in the expression when matching. That means that an expression like CVCVCVCVCV (five alternating consonant-vowel pairs) matches words like "literature" as well as phrase fragments like "have become" or "was used as a".

To restrict where spaces can be placed, use "quoted phrases" in the pattern. Within the quotation marks, spaces will not be inserted unless the pattern specifically allows them. So the expression "CVCVCVCVCV" only matches single 10-letter words; "CVCVC-VCVCV" matches either single words or evenly split pairs of 5-letter words; "CVCVC VCVCV" matches only pairs of 5-letter words.

Quotation marks can be used on all or part of the query, as desired.

& (Intersection)

By analogy with the standard | (alternation) operator, the & operator requires both sides to match for the pattern to match. This is useful for applying several constraints in parallel. For example, "_*a_*&_*b_*&_*c_*&_{5,}" matches single words that contain all of "a", "b" and "c" (but in no particular order) and are also at least 5 letters long overall.

(Nutristatic compiles intersections lazily, so even large numbers of clauses are cheap to parse — though the search itself still gets harder with each added constraint's rarity.)

<Anagrams>

Text inside <angle brackets> matches anything that contains the same parts but in any order. So <act> matches "act", "cat", "atc" and so on.

The parts of an anagram are normally letters, but can themselves be any regular expression (in parentheses). For example, <(ag)(m)(ra)__> matches any 7-letter word or phrase containing "ag", "m", "ra", and any other two letters. Anagrams can also be part of a larger expression, if you have partial information about the order of something, such as <aan>g<amr> which matches "anagram" but not "margana".

Remember, if you want to restrict your anagram to single words, use "quoted phrases".

Note: Long anagrams compile fine here (see the differences), but the search for a long anagram still visits many index nodes and may hit the computation limit — the "Try harder" button is your friend.

Note (from the original Nutrimatic guide): the anagram algorithm isn't perfect -- when using anagrams of wildcards or pieces that aren't single letters, sometimes results will be produced that aren't actually an anagram of the input. If you're using fancy anagram expressions, double-check what you get back.

Examples from MSPH 12(3)

These worked examples are from the original Nutrimatic usage guide (Dan Egnor, using the Microsoft Puzzlehunt 12(3) as a source of wordy puzzles). They all work identically here.

Nutristatic does not solve these puzzles automatically, it only helps with one specific part of the "crank turning". It doesn't help you figure out what to do in the first place, and anyway most of these puzzles require many other steps where a text grep engine is unhelpful.

Camouflage

This puzzle included many lines of text like "MOCHIT HATORY" with a blank in the middle. You were instructed to fill in a letter that would make a word at least 5 letters using a possibly-empty suffix of the first part, a filled in letter, and a possibly-empty prefix of the second part. For that first row, this pattern locates all such words:

"(((((m?o)?c)?h)?i)?t)?_(h(a(t(o(ry?)?)?)?)?)?&_{5,}"

Note that the puzzle directions further require "common non-plural, non-capitalized English" words, and that each letter (A-Z) be used exactly once; these constraints must be enforced by the human solver.

Triple Sec

Among the various other things in this puzzle, you frequently need to rearrange a list of letter triples into a clue with some left over:

MAY SIT TIT BLE COM IKS IAL IMB MON

Six of the triples will be used to make a crossword-type clue; the other three will be left over for future use. The boldface letters indicate the start of a word in the clue (a word break happens just before each boldface letter). This pattern finds the solution:

"<(-may)?(-sit)?(tit)?(ble)?(com)?(iks)?(ial)?(im b)?(-mon)?>"&(_{18})

The anagram operator is used to reorder the triples. Each triple is made individually optional in the anagram, but a constraint is added to the end requiring the answer to be 18 letters long (i.e. consume 6 triples). The expression is quoted; spaces and hyphens are used to indicate where word breaks can occur (boldface letters).

The answer "mayim bialiks sitcom" is among the first results. (Mayim Bialik's sitcom is "Blossom").

Dice

After assembling a cube with spots on it, part of the final step in this puzzle is to treat each side of the cube as a letter bank (such as AEHIMNPRSW) and to determine the phrases that can be formed using that letter bank. We assume the puzzle designer didn't include any extraneous letters, so this is like an anagram with repeated letters collapsed into one. This pattern finds what can be formed from one of these letter banks:

[aehimnprsw]*&_*a_*&_*e_*&_*h_*&_*i_*&_*m_*&_*n_*&_*p_*&_*r_*&_*s_*&_*w_*

This rather cumbersome expression requires a phrase made from letters in the bank and also (this is the cumbersome part) requires each letter to be used at least once somewhere in the phrase. The answer to this one, "new hampshire", comes out on top. (Turns out each of the six sides decodes to a state.)

Hollywood Walk of Fame

One of the steps in the second phase of this puzzle involved identifying a series of rebus-like drawings, and then concatenating the names of the depicted objects with one letter removed from each. For example, one of the stars contained drawings of CHARM, ELTON, CHEST, and ONE. This pattern finds what can be made from this with one letter dropped from each:

(c?h?a?r?m?&_{4})(e?l?t?o?n?&_{4})(c?h?e?s?t?&_{4})(o?n?e?&_{2})

Within each word, every letter is optional, but a constraint is also added to the subpattern requiring that N-1 of the letters be used. The answer "charlton heston" comes out on top.

Murder By Depth

After finding this puzzle in a pool of water, you need to take a number of nonsense strings like "SEPAP" and "HEGM" and "add water" by inserting the letters W, A, T, E, R -- in that order, but not necessarily consecutively -- into the given letters. This pattern finds the answer for HEGM:

<waterhegm>&_*w_*a_*t_*e_*r_*

The word is required to be an anagram of <waterhegm> with the added constraint that the letters in "water" appear in that order. The answer "wheat germ" shows up on top.

Mortal Jeopardy

This meta for the third round was quite involved, but in the end produced a series of (mostly) three-letter chunks. The series was in the right order, but the chunks were scrambled internally:

<het><ral><seg><tan><rut><bla><oody><afl><ndi><cin><awe><ter>

The answer, "the largest natural body of land in ice water", shows up first. (Interpreting it requires some cleverness by the human solver!)

How It Works (and What Makes a Query Slow)

Internally, Nutristatic uses the same data structure as Nutrimatic: a trie of every word and phrase that occurs in Wikipedia at least five times, where every node carries a frequency count. The English index packs about 5.7 billion indexed word positions into a 1.3 GB file, in a format byte-compatible with Nutrimatic's tools.

When searching for a pattern, the engine runs a best-first search through the trie using a priority queue of nodes sorted by frequency. The queue initially contains only the root node. As long as the queue is not empty, the algorithm takes the most frequent node from the queue and examines it for compatibility with the search pattern. If the path to that node matches the pattern, it's printed as output. If the path to that node is a possible prefix for the pattern (not necessarily matching the pattern itself), then the node's children are all added to the queue. Then it returns to the next best node in the queue and continues the process.

Whenever a space is encountered in the search, in addition to continuing to the children of that node, the search algorithm also continues at the root of the trie, allowing two words or phrases to be spliced together if they match the pattern. A heavy frequency penalty is added for this "reset", so that phrases which occur naturally are considered before such "frankenphrases".

The search expression is evaluated against the trie as a deterministic finite state machine. Where Nutrimatic uses the OpenFST library to build the complete machine before searching, Nutristatic determinizes lazily: states of the machine — including the products implied by & and by anagram constraints — are only constructed when the index walk actually reaches them, and are cached for reuse. Compilation is effectively instant for any pattern; complexity is paid only in proportion to what the search actually explores.

What still takes time is the search itself. The search pauses at a step budget and offers "Try harder" to continue with a doubled budget. In the default streaming mode there is a second cost: index pieces are fetched over the network as the search walks them, so a broad first-time query also downloads a few megabytes (with fetched pieces cached for next time; the downloaded-index mode has no such cost at all).

Because the search proceeds forward from the start of the answer string, patterns which are constrained at the beginning work much better than patterns which are constrained at the end. For example, the_* gives results very quickly; _*est takes a long time to return. The first pattern can walk the subtrie starting with "the"; the second pattern effectively walks every word and phrase in the index in frequency order, checking for each one whether it ends in "est".

Adding more constraints always helps the search -- for example, quoting the search string is good if you don't expect the answer to contain multiple words (or you know where word breaks go).


Nutristatic is derived from Nutrimatic (GPL, by Dan Egnor and contributors); the syntax documentation and examples above are adapted from the original Nutrimatic usage guide. Updated Aug 20, 2026.

Impressum · Datenschutz