Asking small decision models (Jev + Laya) about first names
Thousands of contacts without salutation, and a model that has to guess from a first name alone.
TL/DR: I sent 20 hand-picked first names went into three variants of a "System One" decision model. Asked each for female, male, multiple or unclear... Jev scored 18 of 20, local Laya 13, a second local mode landed at 10.
In our Salesforce environment, roughly 6,700 household contacts carry no salutation. Why? Mostly external systems (like sign-up forms) gathered those contacts, but didn't ask for a salutation 🤷 However, every one of them has a first name. And with the recent hype around convenient new classifiers, I wanted to know: can such a small model look at a first name and identify the contact as female, male, several people, or can't tell?
The test setup: two backends, the same prompt used with both, and a small sample test set. Laya runs locally under Unsloth (nothing leaves my machine), while Jev from TypeSafe is a hosted API.

The setup
A small query gets contacts with empty salutations, private households only (because contacts in organisations are a different kind of weird 😅):
select c.id, c.firstname
from raw.salesforce.contact c
join raw.salesforce.account a on a.id = c.accountid
where coalesce(c.salutation, '') = ''
and a.type = 'Household'A stdlib-only Python script does the classifying, one request per contact. The only thing it gets to use is the first name. And the decision model gets a single choice question with four options, female, male, multiple and n/a. Every run produces a CSV holding the choice and the confidence.
Per its instructions, the model assumes a Swiss context (Swiss German, French, Italian and international names). These criteria ended up working best:
female/male: a single person. Compound first names count as one too, whether written as one word or two (Anna Maria, Hanspeter, Jean Daniel)multiple: two or more people, joined by a separator such as&,+,undoretn/a: unisex or ambiguous names, initials only, titles, or anything that isn't a first name
import csv
import json
import os
import urllib.request
from pathlib import Path
HERE = Path(__file__).parent
URL = os.environ.get("DECISION_API_URL", "http://localhost:8888/v1/systemone")
MODEL = os.environ.get("DECISION_MODEL", "laya")
HEADERS = {"Content-Type": "application/json"}
if os.environ.get("DECISION_API_KEY"): # not needed for local
HEADERS["Authorization"] = "Bearer " + os.environ["DECISION_API_KEY"]
QUESTION = {
"type": "choice",
"instructions": (
"The state holds the first-name field of a contact (`firstname`). "
"Classify which salutation fits, judging only from the first name. "
"Context: Switzerland, names may be Swiss, German, French, Italian "
"or international."
),
"criteria": {
"female": (
"A single person with a female first name. Compound first names "
"written as one or two words (Annemarie, Anna Maria) are one "
"person."
),
"male": (
"A single person with a male first name. Compound first names "
"written as one or two words (Hanspeter, Jean Daniel) are one "
"person."
),
"multiple": (
"Two or more people joined by a separator such as '&', '+', "
"'und' or 'et' between their names."
),
"n/a": (
"Unisex or ambiguous name, only an initial, a title, or not a "
"first name."
),
},
}
def classify(firstname: str) -> dict:
body = {
"model": MODEL,
"state": {"firstname": firstname},
"questions": {"salutation": QUESTION},
}
req = urllib.request.Request(
URL,
data=json.dumps(body).encode("utf-8"),
headers=HEADERS,
)
with urllib.request.urlopen(req, timeout=120) as resp: # raises on non-2xx
return json.load(resp)["answers"]["salutation"]
def main() -> None:
sample = HERE / os.environ.get("SAMPLE_FILE", "sample_contacts.json")
contacts = json.loads(sample.read_text("utf-8"))
rows = []
for c in contacts:
a = classify(c["firstname"])
rows.append({
"id": c["id"],
"firstname": c["firstname"],
"expected": c.get("expected", ""),
"choice": a["choice"],
"confidence": round(a["confidence"], 3),
"probabilities": json.dumps(a["probabilities"]),
})
print(f"{c['firstname']:<22} {c.get('expected', ''):<9} "
f"{a['choice']:<9} {a['confidence']:.3f}")
with open(HERE / f"results_{sample.stem}_{MODEL}.csv", "w", newline="", encoding="utf-8") as f:
w = csv.DictWriter(f, fieldnames=list(rows[0]))
w.writeheader()
w.writerows(rows)
if __name__ == "__main__":
main()Let's test!
To illustrate the result, I picked 20 first names from the list of contacts, five per category (male, female, multiple, unclear) and each with an expected label.
| Group (5 names each) | Jev | Laya typed | Laya multilingual |
|---|---|---|---|
| female | 5 | 4 | 5 |
| male | 5 | 3 | 3 |
| multiple | 5 | 5 | 0 |
| unclear | 3 | 1 | 2 |
| total | 18 / 20 | 13 / 20 | 10 / 20 |
- Jev's two misses: Andrea (0.97) and Simone (0.85), both returned as
female. That is a reasonable result in Swiss German. However, in Italian both are male names, which is why I would label themn/a - Swiss first names tripped up both local modes 😅 Typed mode turned Reto and Beat into
femaleand made Rahelmale; multilingual mode returnedn/afor Reto andfemalefor Beat. Jev got all three right - Multilingual mode never picked
multiple. All five pairs came backfemale, four at confidence 1.0... which means its confidence says nothing 😅 - Typed mode nails
multipleevery time. On single names the confidence lands between 0.006 and 0.11, correct and incorrect answers overlap (Walter at 0.057 is right, Simone at 0.088 is wrong).
Jev's confidence, by contrast, appears usable. Anything at 0.85 or above was right, and the single uncertain case (an unusual name, correctly n/a) landed at 0.26. A cutoff around 0.8 would push only edge-case-names into manual review.
All 3 variants were quite fast: A set of 1000 contacts takes about 5 to 10 seconds locally (Laya) after adding a little parallelism to the process. I didn't test the speed of Jev (for the reason below), but others did, so it would presumably land in the same ballpark.
The privacy part
The Jev runs sent first names (nothing else) from my database to a third-party API, currently not offering data residency in the EU. A real run over thousands of contacts would require a data-protection decision first. Laya, on the other hand, never leaves my machine. That is a real argument in its favor, even if it requires further tinkering... or acceptance that 13 of 20 is better than nothing 😅
What I haven't tried yet
Plain code with regex on separators (&, und, +, et, /) could catch multiple on its own, leaving the model only to decide between female / male / n/a for the rest. Another candidate: a hard rule forcing n/a for names that flip between languages (Andrea, Simone, Nicola, Dominique). And separate yes/no questions per label (instead of one choice) might behave differently again.
