FREE · EVERY WEEKDAY

Stay up to date on AI

The most important AI news — straight to your inbox.

Practical guide

Extract structured data with local AI

Turn document text into defined fields. Use a schema for the output, represent missing information explicitly and check every value against the source.

AI News TrendUpdated 8 min read
In this guide

Extraction is useful when the answer needs to become a row, a form or a record. The difficult part is not just producing JSON. It is knowing which event a date describes, which person owns a statement and when a value should be absent.

Define the fields before writing the prompt

Specify what each field means, its data type and how missing information should be represented. A person's name can refer to a writer, a speaker or a project owner. A date can refer to publication, a decision or a planned launch. Choose one meaning for the task.

The five-field contract used in our recorded Swedish examples.
FieldMeaningMissing value
organizationThe organisation behind the event.null
personThe responsible person or speaker, not the article's author.null
dateThe event date, formatted as YYYY-MM-DD.null
amountSekAn explicitly stated SEK amount as a number.null
quoteA verbatim quotation without enclosing quotation marks.null

Include documents with missing dates, multiple people and corrected amounts. Do not reward a filled field when the source contains no support for it. A complete-looking record can be less useful than an honest missing value.

Use a schema when the workflow requires a fixed shape

Ollama accepts a JSON schema in the chat request's format field. Define required keys, allowed types and whether additional properties are allowed. Mention the schema in the instruction too. This controls the output shape; it does not establish the truth of a value.

The following two-field example demonstrates the request structure with fictional English text. It expects organization: "Tallvik Teknik" and amountSek: null. It has not been included in the recorded Swedish results.

Show a minimal JavaScript request
// Node.js 20+. This is a new workflow example, not the recorded pilot.
const schema = {
  type: 'object',
  properties: {
    organization: { type: ['string', 'null'] },
    amountSek: { type: ['integer', 'null'] }
  },
  required: ['organization', 'amountSek'],
  additionalProperties: false
};
const source = 'Tallvik Teknik is planning an internal document trial. '
  + 'The company has not announced a budget.';
const response = await fetch('http://127.0.0.1:11434/api/chat', {
  method: 'POST',
  headers: { 'Content-Type': 'application/json' },
  body: JSON.stringify({
    model: 'qwen3:8b', stream: false, think: false,
    format: schema,
    options: { temperature: 0, num_ctx: 4096, num_predict: 256 },
    messages: [{ role: 'user', content:
      'Extract the organization and an explicitly stated SEK budget. '
      + 'Use null for missing information. Return only JSON matching '
      + JSON.stringify(schema) + '\nSOURCE:\n' + source }]
  })
});
if (!response.ok) throw new Error('HTTP ' + response.status);
const answer = await response.json();
if (answer.done_reason === 'length') throw new Error('Incomplete answer');
const fields = JSON.parse(answer.message.content);
// Validate the keys and types in this two-field schema.
if (!fields || typeof fields !== 'object' || Array.isArray(fields)
  || Object.keys(fields).sort().join(',') !== 'amountSek,organization'
  || !(fields.organization === null || typeof fields.organization === 'string')
  || !(fields.amountSek === null || Number.isInteger(fields.amountSek))) {
  throw new Error('The response does not match the schema');
}
// Check this example against its source: a name is given, a budget is not.
if (fields.organization !== 'Tallvik Teknik' || fields.amountSek !== null) {
  throw new Error('The fields do not match the source');
}
console.log(fields);

Download the specified model first. Save the code as a .mjs file and run it with Node.js. Add a schema validator before accepting the returned object into another system.

Our older Swedish examples used free-text output with a written JSON instruction, without schema enforcement. Keep that distinction when interpreting the examples: they show failures of that configuration, not a comparison of schema-enforced extraction.

Validate format and factual content separately

  1. Check completion. Reject truncated or empty answers rather than treating partial content as a finished record.
  2. Parse and validate. Check all required keys, types, null handling and unexpected properties. A JSON parser alone does not perform schema validation.
  3. Check source support. For each populated field, identify the sentence or span that supports it. Validate the date and amount against the event you intended to extract.
  4. Handle failures explicitly. Retain the source and the failed response for review. Avoid silently filling absent values from assumptions or earlier documents.

For an evaluation set, write an answer key first. Track format failures separately from wrong values and missing supported values. Exact comparison is reproducible but strict: a semantically equivalent spelling can fail. Inspect the original response alongside the score.

If you need evidence in a live workflow, consider adding a source-span field and verify that the quoted span exists in the input. A model-generated explanation is still something to check.

Inspect a recorded Swedish example

These short fictional sources test author versus speaker, absent information and a corrected budget. Source text, expected fields and model output are shown together. The Swedish text is preserved; the controls are in English.

Source and recorded response

Fictional Swedish material · original text preserved · explanations in English

Roles, dates and amounts

Source text · Swedish

Nordhamn Support startar ett pilotprojekt den 14 oktober 2026. Projektledaren Karin Ek ansvarar för projektet. Budgeten är 125 000 kronor. Karin Ek säger: ”Vi börjar med ett avgränsat försök.” Artikeln är skriven av Leo Berg.

Expected fields
{
  "organization": "Nordhamn Support",
  "person": "Karin Ek",
  "date": "2026-10-14",
  "amountSek": 125000,
  "quote": "Vi börjar med ett avgränsat försök."
}

qwen3:1.7b · Run 1

{
  "organization": "Nordhamn Support",
  "person": "Karin Ek",
  "date": "2026-10-14",
  "amountSek": null,
  "quote": "Vi börjar med ett avgränsat försök."
}

Parses as a JSON object · 4/5 expected field values match. This is not a full JSON Schema validation.

qwen3:1.7b · Run 2

{
  "organization": "Nordhamn Support",
  "person": "Karin Ek",
  "date": "2026-10-14",
  "amountSek": null,
  "quote": "Vi börjar med ett avgränsat försök."
}

Parses as a JSON object · 4/5 expected field values match. This is not a full JSON Schema validation.

qwen3:4b · Run 1

Incomplete response. This request reached the output limit of 384 tokens. It returned reasoning rather than a complete final answer despite the request for no thinking. This configuration needs further investigation before its model quality can be judged.

Okay, let's tackle this problem step by step. So, I need to parse the given Swedish text and extract specific information into a JSON object with five keys: organization, person, date, amountSek, and quote. The user wants only the exact values from the text, no calculations, and null if something's missing.

First, let me read through the source text carefully.

The text says: "Nordhamn Support startar ett pilotprojekt den 14 oktober 2026. Projektledaren Karin Ek ansvarar för projektet. Budgeten är 125 000 kronor. Karin Ek säger: ”Vi börjar med ett avgränsat försök.” Artikeln är skriven av Leo Berg."

Let me break this down sentence by sentence.

1. "Nordhamn Support startar ett pilotprojekt den 14 oktober 2026." 
   - This translates to "Nordhamn Support starts a pilot project on October 14, 2026."
   - So the organization here is "Nordhamn Support". The date is October 14, 2026. In YYYY-MM-DD format, that's 2026-10-14.

2. "Projektledaren Karin Ek ansvarar för projektet."
   - "Project leader Karin Ek is responsible for the project."
   - So the person here is Karin Ek.

3. "Budgeten är 125 000 kronor."
   - "The budget is 125,000 kronor."
   - The amount in SEK is 125000 (since it's written as 125 000, which is

Not a valid JSON object · 0/5 expected field values match. This is not a full JSON Schema validation.

qwen3:4b · Run 2

Incomplete response. This request reached the output limit of 384 tokens. It returned reasoning rather than a complete final answer despite the request for no thinking. This configuration needs further investigation before its model quality can be judged.

Okay, let's tackle this problem step by step. So, I need to parse the given Swedish text and extract specific information into a JSON object with five keys: organization, person, date, amountSek, and quote. The instructions say to return only the JSON, no extra info, and use null if something's missing.

First, let me read the source text carefully.

The text says: "Nordhamn Support startar ett pilotprojekt den 14 oktober 2026. Projektledaren Karin Ek ansvarar för projektet. Budgeten är 125 000 kronor. Karin Ek säger: ”Vi börjar med ett avgränsat försök.” Artikeln är skriven av Leo Berg."

Let me break it down sentence by sentence.

1. "Nordhamn Support startar ett pilotprojekt den 14 oktober 2026." 
   - This translates to "Nordhamn Support starts a pilot project on October 14, 2026." So the organization here is Nordhamn Support. The date is October 14, 2026. But the date needs to be in YYYY-MM-DD format. October 14, 2026 would be 2026-10-14.

2. "Projektledaren Karin Ek ansvarar för projektet."
   - "Project leader Karin Ek is responsible for the project." So the person here is Karin Ek.

3. "Budgeten är 125 000 kronor."
   - "The budget is 125,000 kronor." So the amount in SEK is 125000 (since it's written with a space in

Not a valid JSON object · 0/5 expected field values match. This is not a full JSON Schema validation.

qwen3:8b · Run 1

{
  "organization": "Nordhamn Support",
  "person": "Karin Ek",
  "date": "2026-10-14",
  "amountSek": 125000,
  "quote": "Vi börjar med ett avgränsat försök."
}

Parses as a JSON object · 5/5 expected field values match. This is not a full JSON Schema validation.

qwen3:8b · Run 2

{
  "organization": "Nordhamn Support",
  "person": "Karin Ek",
  "date": "2026-10-14",
  "amountSek": 125000,
  "quote": "Vi börjar med ett avgränsat försök."
}

Parses as a JSON object · 5/5 expected field values match. This is not a full JSON Schema validation.

Missing information should be null

Source text · Swedish

Tallvik Teknik förbereder ett internt försök med dokumentanalys. Bolaget har inte angett datum, budget eller vem som ansvarar. Texten innehåller inga direktcitat.

Expected fields
{
  "organization": "Tallvik Teknik",
  "person": null,
  "date": null,
  "amountSek": null,
  "quote": null
}

qwen3:1.7b · Run 1

{
  "organization": null,
  "person": null,
  "date": null,
  "amountSek": null,
  "quote": null
}

Parses as a JSON object · 4/5 expected field values match. This is not a full JSON Schema validation.

qwen3:1.7b · Run 2

{
  "organization": null,
  "person": null,
  "date": null,
  "amountSek": null,
  "quote": null
}

Parses as a JSON object · 4/5 expected field values match. This is not a full JSON Schema validation.

qwen3:4b · Run 1

Incomplete response. This request reached the output limit of 384 tokens. It returned reasoning rather than a complete final answer despite the request for no thinking. This configuration needs further investigation before its model quality can be judged.

Okay, let's tackle this problem step by step. So, I need to extract specific information from the given Swedish text and return a JSON object with five keys: organization, person, date, amountSek, and quote. The instructions say to use null if the information isn't present. Let me read through the text carefully.

The text says: "Tallvik Teknik förbereder ett internt försök med dokumentanalys. Bolaget har inte angett datum, budget eller vem som ansvarar. Texten innehåller inga direktcitat."

First, I'll break down each part.

1. **Organization (organisationen bakom händelsen)**: The text mentions "Tallvik Teknik" as the company preparing an internal document analysis trial. So the organization here is "Tallvik Teknik".

2. **Person (personen som uttalar sig eller ansvarar för händelsen)**: The text states that the company hasn't specified who is responsible. So there's no person mentioned here. The answer should be null.

3. **Date (datumet för händelsen i YYYY-MM-DD)**: The text says "Bolaget har inte angett datum", which translates to "The company has not specified a date". So date is null.

4. **amountSek (det uttryckligen angivna beloppet i kronor som ett tal)**: The text mentions "budget" but says the company hasn't specified it. So there's no amount given here. amountSek should be null.

5. **quote (ett exakt direktcitat inklusive originalets ord men utan omgivande citationstecken)**: The text states "Texten innehåller inga direktcitat." which means "The text contains no direct quotes." So there's

Not a valid JSON object · 0/5 expected field values match. This is not a full JSON Schema validation.

qwen3:4b · Run 2

Incomplete response. This request reached the output limit of 384 tokens. It returned reasoning rather than a complete final answer despite the request for no thinking. This configuration needs further investigation before its model quality can be judged.

Okay, let's tackle this problem. So, I need to extract specific information from the given Swedish text and return a JSON object with five keys: organization, person, date, amountSek, and quote. The user said to use null if the information isn't present.

First, let me read the source text carefully: "Tallvik Teknik förbereder ett internt försök med dokumentanalys. Bolaget har inte angett datum, budget eller vem som ansvarar. Texten innehåller inga direktcitat."

Breaking it down sentence by sentence.

First sentence: "Tallvik Teknik förbereder ett internt försök med dokumentanalys." Translates to "Tallvik Teknik is preparing an internal document analysis trial." So here, the organization is Tallvik Teknik.

Second sentence: "Bolaget har inte angett datum, budget eller vem som ansvarar." Which means "The company has not specified date, budget, or who is responsible." So here, the company (organization) hasn't given a date, budget, or responsible person.

Third sentence: "Texten innehåller inga direktcitat." Translates to "The text contains no direct quotes." So there are no quotes here.

Now, the required keys:

1. organization: The organization behind the event. From the text, it's Tallvik Teknik. So that's straightforward.

2. person: The person who speaks or is responsible for the event. The text says "Bolaget har inte angett vem som ansvarar" (the company hasn. specified who is responsible). So there's no person mentioned here. So person should be null.

3. date: The date of the event in YYYY-MM-DD. The text says "Bolaget har inte angett datum" (the company has not

Not a valid JSON object · 0/5 expected field values match. This is not a full JSON Schema validation.

qwen3:8b · Run 1

{
  "organization": "Tallvik Teknik",
  "person": null,
  "date": null,
  "amountSek": null,
  "quote": null
}

Parses as a JSON object · 5/5 expected field values match. This is not a full JSON Schema validation.

qwen3:8b · Run 2

{
  "organization": "Tallvik Teknik",
  "person": null,
  "date": null,
  "amountSek": null,
  "quote": null
}

Parses as a JSON object · 5/5 expected field values match. This is not a full JSON Schema validation.

A corrected amount and a Swedish date

Source text · Swedish

Rättelse: Lindåra kommun fattade beslut om ett bibliotekstest den 3 oktober 2026. Ansvarig samordnare är Amir Lund. Budgeten är 350 000 kronor, inte de 3 500 000 kronor som angavs i en äldre text. Amir Lund säger: ”Vi rättar uppgiften och gör om kontrollen.”

Expected fields
{
  "organization": "Lindåra kommun",
  "person": "Amir Lund",
  "date": "2026-10-03",
  "amountSek": 350000,
  "quote": "Vi rättar uppgiften och gör om kontrollen."
}

qwen3:1.7b · Run 1

{
  "organization": "Lindåra kommun",
  "person": "Amir Lund",
  "date": "2026-10-03",
  "amountSek": null,
  "quote": "Vi rättar uppgiften och gör om kontrollen."
}

Parses as a JSON object · 4/5 expected field values match. This is not a full JSON Schema validation.

qwen3:1.7b · Run 2

{
  "organization": "Lindåra kommun",
  "person": "Amir Lund",
  "date": "2026-10-03",
  "amountSek": null,
  "quote": "Vi rättar uppgiften och gör om kontrollen."
}

Parses as a JSON object · 4/5 expected field values match. This is not a full JSON Schema validation.

qwen3:4b · Run 1

Incomplete response. This request reached the output limit of 384 tokens. It returned reasoning rather than a complete final answer despite the request for no thinking. This configuration needs further investigation before its model quality can be judged.

Okay, let's tackle this problem step by step. So, I need to parse the given Swedish text and extract five specific pieces of information into a JSON object. The keys are organization, person, date, amountSek, and quote. If any of these aren't present, I should use null.

First, let me read through the source text carefully.

The text says: "Rättelse: Lindåra kommun fattade beslut om ett bibliotekstest den 3 oktober 2026. Ansvarig samordnare är Amir Lund. Budgeten är 350 000 kronor, inte de 3 500 000 kronor som angavs i en äldre text. Amir Lund säger: ”Vi rättar uppgiften och gör om kontrollen.”"

Okay, breaking it down:

- "Rättelse" means correction, so this is a correction notice.
- "Lindåra kommun fattade beslut om ett bibliotekstest den 3 oktober 2026." translates to "Lindåra Municipality made a decision about a library test on October 3, 2026."
- "Ansvarig samordnare är Amir Lund." means "The responsible coordinator is Amir Lund."
- "Budgeten är 350 000 kronor, inte de 3 500 000 kronor som angavs i en äldre text." translates to "The budget is 350,000 kronor, not the 3,500,000 kronor that was stated in an older text."
- "Amir Lund säger: ”Vi rättar uppgiften och gör om kontrollen.”" means "Am

Not a valid JSON object · 0/5 expected field values match. This is not a full JSON Schema validation.

qwen3:4b · Run 2

Incomplete response. This request reached the output limit of 384 tokens. It returned reasoning rather than a complete final answer despite the request for no thinking. This configuration needs further investigation before its model quality can be judged.

Okay, let's tackle this problem step by step. So, I need to parse the given Swedish text and extract specific information into a JSON object with five keys: organization, person, date, amountSek, and quote. The user specified that I should return only a JSON object with those keys, using null if something is missing. Also, I shouldn't add extra info or calculate amounts that aren't explicitly stated.

First, let me read through the source text carefully.

The text says: "Rättelse: Lindåra kommun fattade beslut om ett bibliotekstest den 3 oktober 2026. Ansvarig samordnare är Amir Lund. Budgeten är 350 000 kronor, inte de 3 500 000 kronor som angavs i en äldre text. Amir Lund säger: ”Vi rättar uppgiften och gör om kontrollen.”"

Okay, breaking it down.

1. **organization**: The organization behind the event. The text mentions "Lindåra kommun", which translates to Lindåra Municipality. So the organization here is Lindåra kommun.

2. **person**: The person who is speaking or responsible for the event, not the article's author. The text says "Ansvarig samordnare är Amir Lund." So the person here is Amir Lund.

3. **date**: The date of the event in YYYY-MM-DD format. The text states "den 3 oktober 2026". So October 3, 2026. In YYYY-MM-DD, that's 2026-10-03.

4. **amountSek**: The explicitly stated amount in SEK as a number. The text mentions "Budgeten är 350 0

Not a valid JSON object · 0/5 expected field values match. This is not a full JSON Schema validation.

qwen3:8b · Run 1

{
  "organization": "Lindåra kommun",
  "person": "Amir Lund",
  "date": "2026-10-03",
  "amountSek": 350000,
  "quote": "Vi rättar uppgiften och gör om kontrollen."
}

Parses as a JSON object · 5/5 expected field values match. This is not a full JSON Schema validation.

qwen3:8b · Run 2

{
  "organization": "Lindåra kommun",
  "person": "Amir Lund",
  "date": "2026-10-03",
  "amountSek": 350000,
  "quote": "Vi rättar uppgiften och gör om kontrollen."
}

Parses as a JSON object · 5/5 expected field values match. This is not a full JSON Schema validation.

Recorded local responses. The source material describes invented events.

Choose “Missing information should be null” and qwen3:1.7b. The response parses as JSON, but the organisation is also returned as null even though the text names it. In other examples, the small model omits a required field. Both are useful reminders to check beyond successful parsing.

qwen3:8b matched all expected fields across these three texts and two runs per text. That is an observation about these examples, not a reliability estimate for invoices or other unseen documents.

Method and reproduction

The examples were recorded locally using Ollama 0.18.2 on 2026-10-05. Each model received the same instruction and source in a fresh conversation. There were three texts for each task and two runs per text. The three models belong to one family.

Requests used temperature 0.2, context 4096 tokens, an output limit of 384 tokens and seeds 42 and 43. We requested think: false. No system prompt or schema-enforced output was used.

The extraction check first parses an object, then compares the five expected field values exactly with the answer key. Missing keys and wrong types fail the corresponding comparison. Extra keys do not earn credit and are not penalised by this field count; validate them separately in a real schema.

The local qwen3:4b configuration reached the output limit in every recorded request and did not produce a complete final answer. The evidence includes those failures. Different model templates, versions or output limits need a separate test; these examples do not establish that the model itself is generally unsuitable.

Original Swedish instruction and model identities

Läs källtexten och returnera endast ett JSON-objekt med dessa fem nycklar: organization (organisationen bakom händelsen), person (personen som uttalar sig eller ansvarar för händelsen, inte artikelns författare), date (datumet för händelsen i YYYY-MM-DD), amountSek (det uttryckligen angivna beloppet i kronor som ett tal), quote (ett exakt direktcitat inklusive originalets ord men utan omgivande citationstecken). Använd null när uppgiften saknas. Lägg inte till information och räkna inte ut belopp som inte uttryckligen anges.

  • qwen3:1.7b · Q4_K_M
    8f68893c685c3ddff2aa3fffce2aa60a30bb2da65ca488b61fff134a4d1730e7
  • qwen3:4b · Q4_K_M
    359d7dd4bcdab3d86b87d73ac27966f4dbb9f5efdfcc75d34a8764a09474fae7
  • qwen3:8b · Q4_K_M
    500a1f067a9f782620b40bee6f7b0c89e17ae61f686b92c24933e4ca4b2b8b41
Repeat the examples on your computer
  1. Install Ollama and Node.js 20 or later. Save the dataset and runner in the same folder, keeping their file names.
  2. Download the three listed model tags if they are not already present. Downloads need internet access and disk space.
  3. Open PowerShell in that folder and run:
ollama pull qwen3:1.7b
ollama pull qwen3:4b
ollama pull qwen3:8b
node .\run-swedish-tests.mjs

The runner makes local API requests and saves swedish-benchmark.json beside the script. That file is overwritten if it already exists. Rename previous results before starting another run.

It downloads no models and changes no global runtime settings. It stops if a model outside the example set is already loaded, to avoid replacing someone else's workload. Use an otherwise idle runtime.

The runner records model identities, instructions and answers. It does not collect hardware specifications, system build numbers, file paths or performance measurements. Compare your answers and settings with the example data; software and model updates can change the result.

Documentation

Updated 5 October 2026. Schema-enforced extraction is a suggested workflow here; it has not been measured in the supporting Swedish pilot.

Continue with a related guide

Get startedRun local AI on Windows with OllamaSummarize documentsChoose a local model for document summarization

These guides combine practical instructions, linked documentation and clearly labelled local examples. New model releases do not become recommendations until the relevant task has been checked.

Our editorial approach and corrections →