Skip to main content

Data import and demo data

Use the shipped skill

Prefer the author-import-package skill (in your scaffolded repo's .claude/skills/) over hand-rolling this.

Kandra has one subsystem for getting data into a configuration: an import package (a JSON file) applied by an idempotent import engine. The same format and the same engine serve two jobs:

  • Demo data. Your configuration ships a demo company as a package with "origin": "demoSeed". A demo installation applies it and gets a populated database.
  • Real import. An administrator, an endpoint call or your own data processor applies a package with "origin": "import": master data from an old system, a price list from a supplier, a nightly feed.

Every row the engine writes goes through the real CRUD services: behaviors, validators, numbering, posting and authorization all run exactly as if a user had entered the data by hand. There is no back door into the tables.

The Import package format reference page lists every field and rule. This page explains how the pieces fit together and walks through real examples.

Concepts​

What "idempotent" means here​

You can apply the same package any number of times. The second run changes nothing; a run after you edit one object in the package changes that object and nothing else. This works because:

  • every object in a package has a $id, unique within the package;
  • the package has an id of its own ("package": "acme-demo");
  • every row the engine creates is stamped with that pair, plus a hash of the object as written.

On the next run the engine finds the row by package id and $id, compares the hash, and reports one outcome per object: Inserted, Unchanged, Updated, or one of the Skipped* outcomes when a user has edited or deleted the row since. What it does in each case is set by per-step policies; the defaults protect user edits (see Policies).

There is no global transaction. Each object is written in its own unit of work, as one HTTP request would be, and a failed run can simply be run again: it continues where it stopped.

The import stamp​

The stamp lives on the row itself, in the change-tracking columns every entity already has (Import_* columns, part of ChangesInfo):

FieldMeaning
OriginImport or DemoSeed (from the package header)
SourceThe package id
KeyThe object's $id
SourceHashSHA-256 of the object as written in the package
ImportedAtWhen the engine last wrote the row

Because the stamp is part of the row, it survives a soft delete and disappears with a hard one. "Changed by a user after import" is simply Modified > ImportedAt, and an ordinary save from the UI keeps the stamp. There is no separate ledger table to keep in sync.

Where you see it in the UI​

  • A record's View page, Change Info tab: a Data import block with the origin (Import / Demo data), the package, the key, when it was imported and Changed after import.
  • Change history tables (and a constant's history): an Origin column. Every change is marked Interactive, API key, MCP agent, Scheduled job, Import, Demo data or System. For Import and Demo data rows the chip links to that import run's report (administrators only).
  • Admin → Data import (/admin/import): the embedded packages with their version and the version last applied, a dry run and run button, the report of each run and the run history. A report's Changes made by this run button lists every change event the run produced.

Walkthrough: extending the Kandra WMS demo package​

The reference configuration ships its demo company as KandraWms.Application/ImportPackages/kandrawms-demo-uk.json (27 steps: units of measure, a branch, warehouses, counterparties, items, then documents). Your scaffolded configuration has the same slot, Acme.Application/ImportPackages/acme-demo.json, with no steps yet. The steps below use the WMS package because it has data in it. They apply unchanged to yours.

1. Add a warehouse and a goods receipt​

Steps run in order, so a new object goes after anything it references. Add a warehouse to the existing Warehouses step:

{ "dictionary": "Warehouses", "items": [
{ "$id": "wh-main", "code": "WH-MAIN", "name": "Головний склад", "branchId": { "$ref": "br-kyiv01" } },
{ "$id": "wh-lviv", "code": "WH-LVIV", "name": "Склад Львів", "branchId": { "$ref": "br-kyiv01" } }
] }

and a receipt into it in the GoodsReceipts step:

{ "$id": "gr-demo-11", "code": "GR-DEMO-11", "date": "$today-5d",
"counterpartyId": { "$ref": "cp-sup-roshen" },
"toWarehouseId": { "$ref": "wh-lviv" },
"lines": [ { "itemId": { "$ref": "itm-sugar" }, "quantity": 40, "price": 28 } ] }

What to notice:

  • Names are the catalog names (Warehouses, GoodsReceipts: the Name of each Dto's [Kandra*Form] attribute) and properties are the Dto's wire names (camelCase, as in the API). Open the generated JSON Schema in your editor and you get completion for both (see JSON Schema).
  • { "$ref": "wh-lviv" } points at an object defined earlier in the same package. To point at a row that already exists in the database, use { "$code": "WH-MAIN" } instead.
  • "$today-5d" is a date token. It resolves when the package is applied, so the demo data always looks recent.
  • The step has "submit": true, so the receipt is inserted submitted and posts like one a user submitted.

Bump the package's "version" when you change its content. A demo installation applies the new version after an update (see Demo data). The version doesn't affect how a manual run behaves.

2. Dry run​

Open Admin → Data import, pick kandrawms-demo-uk and press Dry run. A dry run parses, validates, resolves every reference, looks up every row and runs the licence check, then writes nothing. On a database that already has the previous version you get:

OutcomeCountWhy
Unchangedeverything elseSame hash as last time
Inserted2wh-lviv and gr-demo-11 are new

A typo shows up here instead of halfway through a real run, as a diagnostic with a JSON path: steps[13].items[10].toWarehouseId: Unknown or later-defined "$ref" "wh-lvov".

3. Run, then run again​

Import is enabled once a dry run of the same source is clean. The real run reports the same two inserts; the new warehouse and receipt now exist and the receipt has posted. Run it once more and every object is Unchanged. The second run writes nothing, so it produces no change events either.

4. Edit an object, run again​

Rename the warehouse in the package ("name": "Склад Львів (центр)") and change the receipt's quantity to 50, then run:

KeyOutcomeWhy
wh-lvivUpdatedDictionary steps default to "onChange": "update"
gr-demo-11SkippedDriftDocument steps default to "onChange": "skip": a changed package document is reported, not rewritten

To make document edits apply, set "onChange": "update" on that step. Updating a submitted document unsubmits it, saves the new content and submits it again, so its register movements and postings are recalculated (see Gotchas).

Now edit the warehouse in the UI and run the package again after another package change: the outcome is SkippedUserModified. The default "onUserModified": "skip" keeps the user's edit. The same applies to a row a user deleted (SkippedDeleted, policy onDeleted).

Entry points​

The admin page​

Admin → Data import (/admin/import, administrators only) lists the packages embedded in your configuration and also takes an uploaded file. The file is uploaded once and both the dry run and the real run use that upload. Each run's report shows counts per outcome as filter chips, a grid of objects, and the diagnostics. /admin/import?run=<run id> opens a run's report directly.

Endpoints​

The same operations are on api/v1/Import, restricted to the admin role. Runs are asynchronous: POST .../run returns 202 with a run id at once, and you poll for the report.

# Packages embedded in the configuration, with the version last applied
curl -H "X-Api-Key: kdr_..." https://acme.example.com/api/v1/Import/packages

# Dry run of an embedded package
curl -X POST -H "X-Api-Key: kdr_..." -H "Content-Type: application/json" \
-d '{"package":"acme-demo","dryRun":true}' https://acme.example.com/api/v1/Import/run

# Run an uploaded file (multipart: file + dryRun + continueOnError)
curl -X POST -H "X-Api-Key: kdr_..." -F file=@prices.json -F dryRun=false -F continueOnError=true \
https://acme.example.com/api/v1/Import/run

# Status and report of one run; recent runs of one package
curl -H "X-Api-Key: kdr_..." https://acme.example.com/api/v1/Import/runs/01a0f801-5869-70e4-80c3-fb95e8f47b14
curl -H "X-Api-Key: kdr_..." "https://acme.example.com/api/v1/Import/runs?package=acme-demo"

The API key must belong to an administrator (see the API key section of Identity & Auth). A second run while one is in progress gets 409. GET .../schema returns the JSON Schema of your configuration's packages. An upload is capped by Import:MaxPackageSizeBytes (default 50 MB).

From code​

IDataImportService (namespace Kandra.DataImport) is the engine itself. It is constructor-injectable and runs as the calling user, synchronously, in the caller's request:

var report = await import.ImportAsync(stream, new ImportOptions(DryRun: false, ContinueOnError: true), cancellationToken);
if (report.HasErrors) { /* report.Diagnostics, report.Items.Where(i => i.Outcome == ImportItemOutcome.Failed) */ }
var inserted = report.Counts.GetValueOrDefault(ImportItemOutcome.Inserted);

ImportOptions has three settings: DryRun, ContinueOnError (when false, the default, the run stops at the first failed object) and RunId (normally left empty). IImportRunService.Start(...) runs the same engine in the background as the starting user and returns a run id at once; that is what the admin page and the endpoints use. For a real (non-dry) run, both paths record the run in the package's history with its report, so runs started from your own code show up on the admin page too.

Automated import is your own data processor​

The platform ships no import job and no file-format converters. Automating an import is a data processor your configuration owns: it gets the data, turns it into a package if needed, and calls IDataImportService. The processor decides everything that depends on your data: where the file comes from, the size limit, whether to continue past errors, whether to offer a dry run. Writing a package and running the engine are the platform's job.

Worked example: items from a CSV file​

The reference configuration has ImportItemsCsv, which takes a CSV file with the header code;name;unit;parent and imports the rows as items. It is a regular data processor: an input Dto, a result Dto, a validator and a behavior.

// Acme.Forms/DataProcessors/ImportItemsCsv.cs
[KandraDataProcessorForm(Name = "ImportItemsCsv", ValidatorType = typeof(ImportItemsCsvValidator), ResultDtoType = typeof(ImportItemsCsvResult))]
[Description("Imports items from a CSV file (header row: code;name;unit;parent).")]
public class ImportItemsCsvDto : IDataProcessorInputDto
{
[Caption("ImportItemsCsv_File")]
[FilePicker(Extensions = "*.csv")]
public Guid? FileId { get; set; }

[Caption("Import_DryRun")]
public bool DryRun { get; set; }
}

public class ImportItemsCsvResult : IDataProcessorResultDto
{
public Guid RunId { get; set; } // no [Caption]: not shown

[Order(1)] [Caption("Import_Outcome_Inserted")] public int Inserted { get; set; }
[Order(2)] [Caption("Import_Outcome_Updated")] public int Updated { get; set; }
[Order(3)] [Caption("Import_Outcome_Unchanged")] public int Unchanged { get; set; }
[Order(4)] [Caption("ImportItemsCsv_Skipped")] public int Skipped { get; set; }
[Order(5)] [Caption("Import_Outcome_Failed")] public int Failed { get; set; }
[Order(6)] [Caption("Import_Message")] public string? Message { get; set; }
}

The Import_* captions are the engine's own localization keys (the admin page uses them), so only the two ImportItemsCsv_* keys and the processor name need entries in your resx files.

The file field is a Guid? with [FilePicker], as on any form (see File storage). On the processor page the picker uploads the picked file when the user presses Execute, and FileId holds the reference id by the time your behavior runs. Don't make it required in the validator: the form validates before the upload, while FileId is still empty, so a required rule would always fail. Check it in the behavior instead.

// Acme.Application/DataProcessors/ImportItemsCsv.cs (excerpt)
[KandraDataProcessorBehavior(typeof(ImportItemsCsvDto))]
public class ImportItemsCsvBehavior(IBlobService blobs, IDataImportService import, IStringLocalizer localizer)
: IDataProcessorBehavior<ImportItemsCsvDto, ImportItemsCsvResult>
{
public const string PackageId = "acme-items-csv";
private const int MaxFileChars = 5 * 1024 * 1024;

public string ObjectName => "ImportItemsCsv";

public ValueTask OnNewAsync(ImportItemsCsvDto input, CancellationToken cancellationToken) => ValueTask.CompletedTask;

public async Task<ImportItemsCsvResult> OnProcessAsync(ImportItemsCsvDto input, CancellationToken cancellationToken)
{
if (input.FileId is not { } fileId)
{
throw new InvalidStateException(localizer["ImportItemsCsv_NoFile"]);
}

string text;
var (content, _, _) = await blobs.OpenReadAsync(fileId, cancellationToken);
await using (content)
{
using var reader = new StreamReader(content, Encoding.UTF8, detectEncodingFromByteOrderMarks: true);
text = await reader.ReadToEndAsync(cancellationToken);
}

// The processor sets its own limits: the engine accepts whatever package it is given.
if (text.Length > MaxFileChars)
{
throw new InvalidStateException(localizer["ImportItemsCsv_TooLarge"]);
}

var package = ToPackage(text); // CSV -> kandra-import/1, below
using var stream = new MemoryStream(JsonSerializer.SerializeToUtf8Bytes(package));
var report = await import.ImportAsync(stream, new ImportOptions(DryRun: input.DryRun, ContinueOnError: true), cancellationToken);

var counts = report.Counts;
int Count(params ImportItemOutcome[] outcomes) => outcomes.Sum(o => counts.GetValueOrDefault(o));
return new ImportItemsCsvResult
{
RunId = report.RunId,
Inserted = Count(ImportItemOutcome.Inserted, ImportItemOutcome.Adopted),
Updated = Count(ImportItemOutcome.Updated, ImportItemOutcome.Restored),
Unchanged = Count(ImportItemOutcome.Unchanged),
Skipped = Count(ImportItemOutcome.SkippedUserModified, ImportItemOutcome.SkippedDrift, ImportItemOutcome.SkippedDeleted),
Failed = Count(ImportItemOutcome.Failed),
Message = report.Diagnostics.FirstOrDefault(d => d.Severity == ImportDiagnosticSeverity.Error)?.Message
?? report.Items.FirstOrDefault(i => i.Outcome == ImportItemOutcome.Failed)?.Message,
};
}
}

The conversion is ordinary System.Text.Json.Nodes code. Each CSV row becomes one object of an Items step:

var item = new JsonObject
{
["$id"] = "item:" + code, // stable: derived from the business key, not the row number
["isFolder"] = false,
["code"] = code,
["name"] = name,
["category"] = "goods",
["unitOfMeasureId"] = new JsonObject { ["$code"] = unit },
};
if (parent.Length > 0)
{
item["parentId"] = new JsonObject { ["$code"] = parent }; // an existing item folder, by code
}

var package = new JsonObject
{
["format"] = "kandra-import/1",
["package"] = PackageId,
["version"] = 1,
["origin"] = "import",
["configuration"] = "Acme",
["steps"] = new JsonArray
{
new JsonObject { ["dictionary"] = "Items", ["match"] = new JsonArray("code"), ["items"] = items },
},
};

The design decisions behind those few lines:

  • The $id comes from the item code, so the same row in next week's file is the same object, and a changed name comes back as Updated. A $id built from the line number would turn every reordered file into a mess of updates.
  • "match": ["code"] lets the first run adopt an item a user already entered by hand (outcome Adopted) instead of failing on a duplicate code. Match only considers rows without an import stamp.
  • $code lookups for the unit and the parent folder mean the file speaks the user's language (codes), not database ids. The referenced rows must already exist.
  • ContinueOnError: true imports the good rows and reports the bad ones. For a file where a partial import is worse than none, use false together with a dry run first.

Running it on a database that has the folder GOODS-FOOD and the unit PCS: the first run reported Inserted: 2, the second Unchanged: 2. The tests in the reference configuration also cover an adopted row, and an edited row coming back as Updated with its parent set by $code. Both runs appeared in the Run history on Admin → Data import, with their reports. A dry run from your processor returns its report to you but isn't recorded in the history.

The processor gets everything else for free: its page, an ExecuteImportItemsCsv permission, and a place in the job scheduler's processor list.

Scheduling it​

A processor runs on a schedule as a scheduled job: a Cron or Simple (interval) trigger, run as a chosen Run As user. A scheduled run can't ask for input. The runner calls the processor's New step and executes the result, so every input must have a usable default. A file picker has none: a scheduled run of ImportItemsCsv gets FileId == null and fails with "Choose a CSV file".

A watched-folder trigger, which claims files dropped into a folder and hands them to the processor, is not available yet. Until then, schedule processors that fetch their data themselves (the next section), and run file-based ones by hand.

A URL source on a Cron trigger​

A feed published at a URL (a supplier's price list, a partner's stock levels) fits a Cron trigger: the processor needs no input. The shape, as a sketch (it is not part of the reference configuration):

public async Task<SupplierPricesResult> OnProcessAsync(SupplierPricesDto input, CancellationToken cancellationToken)
{
var options = feedOptions.Value; // IOptions<SupplierFeedOptions>, bound from configuration
if (options.Url.Scheme != Uri.UriSchemeHttps)
{
throw new InvalidStateException("The supplier feed must be an https:// URL.");
}

using var request = new HttpRequestMessage(HttpMethod.Get, options.Url);
request.Headers.Authorization = new AuthenticationHeaderValue("Bearer", options.Token);

var lastEtag = (await constants.GetAsync<SupplierFeedEtagConstant>(null, cancellationToken)).Value;
if (!string.IsNullOrEmpty(lastEtag))
{
request.Headers.IfNoneMatch.Add(EntityTagHeaderValue.Parse(lastEtag));
}

using var response = await http.SendAsync(request, cancellationToken); // IHttpClientFactory client
if (response.StatusCode == HttpStatusCode.NotModified)
{
return new SupplierPricesResult { NotModified = true };
}

response.EnsureSuccessStatusCode();
var package = ToPackage(await response.Content.ReadAsStringAsync(cancellationToken));
// ... ImportAsync as in the CSV example ...

if (response.Headers.ETag is { } etag)
{
await constants.SetAsync<string>("SupplierFeedEtag", etag.ToString(), FixedDate, cancellationToken);
}
// ...
}
  • HTTPS only. Refuse anything else in the processor, so a misconfiguration can't send the token in clear text.
  • Keep the secret out of source and out of the database. Bind URL and token from configuration (SupplierFeed__Token as an environment variable, user secrets in development), as you would the JWT secret (Deployment). Don't put it on the input Dto: a scheduled run only gets the defaults from New, and every input field is sent to the browser and shown on the processor page.
  • Skipping unchanged content is an optimization, not a correctness fix. Re-importing the same feed only produces Unchanged outcomes, because the engine is idempotent anyway. An ETag (or a SHA-256 of the body when the server sends no ETag) saves the download and the run. A constant is the natural place for that one value. Write it at a fixed effective date so each write corrects the same row instead of adding history (constants here is IConstantsManageService). Writes are permission-checked, so the job's Run As user needs the constants permission.

Demo data​

A configuration's demo company is an embedded package with "origin": "demoSeed". The scaffolded configuration already has everything wired:

<!-- Acme.Application.csproj -->
<EmbeddedResource Include="ImportPackages\*.json" Exclude="ImportPackages\*.schema.json" />
// Acme.Application/ConfigureServices.cs
services.AddKandraDataImport("Acme");
services.AddImportPackages(typeof(ConfigureServices).Assembly); // every ImportPackages/*.json with format kandra-import/1
services.AddSingleton<IDemoUserProvider, AcmeDemoUsers>();
services.AddKandraDemoSeeding<DomainUser, DomainRole>();

Fill ImportPackages/acme-demo.json with steps; the admin page lists it straight away.

Demo users​

A regular installation creates admin and the two service users (mcp-agent, jobs-scheduler) and no sample users. Sample users exist only on a demo installation. They come from your IDemoUserProvider, not from the package, because a user needs a password and has no change-tracking row to stamp:

internal sealed class AcmeDemoUsers : IDemoUserProvider
{
public IReadOnlyList<DemoUser> GetDemoUsers() =>
[
new DemoUser("user0", "user0@example.com", "User!0", ["user"]),
new DemoUser("user1", "user1@example.com", "User!1", ["user"]),
];
}

Roles must already exist (the seeder creates admin, employee and user).

Free-tier limits​

A demo has to fit the free tier, and the engine checks it before writing anything:

  • Users: at most five in total including admin (the service users don't count). Demo seeding counts existing users plus your demo users and refuses to go over, so keep the list at four or fewer.
  • Documents per month and lines per document: the import engine's licence pre-check counts the documents a package would create. Every one of them is created this month, whatever its date token says. A package over the limit fails its dry run with an error instead of stopping halfway.

How a demo installation is seeded and kept current​

IDemoSeeding (registered by AddKandraDemoSeeding) does the work:

  • SeedAsync(packageId) applies the demo package (by default the first embedded package with origin demoSeed), creates the demo users and writes a marker (DemoInstallation: package and date). Once the marker exists it does nothing, so a demo user that someone deleted is never recreated. A package with errors writes no marker and no users, so calling it again continues the work.
  • GetInstallationAsync() reads the marker: null means this is not a demo installation.
  • ApplyPendingUpdateAsync() runs at every application start. On a demo installation whose embedded demo package has a higher version than the one applied, it applies the package again. The default policies insert what is new and keep what users changed. This is why you bump the version whenever you change the demo package. A failure here is logged and never blocks startup.

The first-run setup wizard (EULA, administrator password, "real company or demo") and an unattended demo mode for hosted demos are not available yet. Both will call SeedAsync.

Gotchas​

  • $ids are forever. The stamp ties a row to its package id and $id. Rename an $id and the engine sees a new object (a second row, or a duplicate-code failure) while the old row stays behind. Renaming the package id has the same effect on every row.
  • Updating submitted documents re-posts them. With "onChange": "update", a changed document is unsubmitted, updated and submitted again in one save. Inventory costing and accounting postings are recalculated from the new content, and a later document that depended on the old quantities can fail its own checks. That is why document steps default to skip. A document made by a createFrom step is never updated by a re-run.
  • $code needs a unique, existing code. The lookup goes to the table the field references, with the entity known from field metadata. It fails if no row has that code. Hierarchical dictionaries can be referenced by folder code too (parentId). An object inserted earlier in the same run can also be found by $code.
  • Any string starting with $ is a token. Only $now, $today and $today±Nd are valid; a literal value such as "$100" is rejected.
  • Date tokens use the caller's time zone. $today is the local date of whoever runs the import: the administrator's browser time zone on the admin page, the job's time zone on a scheduled run. DateTime fields get local midnight converted to UTC, DateOnly fields the local date. Tokens are hashed unresolved, so a package full of $today-5d doesn't count as changed every day.
  • The licence pre-check fails early. A run that would exceed documents per month or lines per document fails as a whole, with a diagnostic at path $, before anything is written.
  • One import at a time per installation. A second run while one is in progress fails (409 from the endpoints). This includes calling IDataImportService from a processor that a package runs in a dataProcessor step: that nested call fails, so don't build that loop.
  • Rights are the caller's. The engine authorizes every write as the user running it. On the admin page that is an administrator; from your processor it is whoever presses Execute; on a scheduled job it is the Run As user. A Run As user without insert rights on items gets a Failed outcome on every row, not a silent skip.
  • Required properties must be written out. A Dto property marked [Required] is an error if the package omits it, even when the Dto would default it. Write explicit values (for example "isIncome": true on a bank receipt). A document's code and isActive are the exceptions: numbering supplies the code, and the step's submit flag sets isActive.
  • Folders before leaves. In a hierarchical dictionary step, a folder must come before the items that reference it with $ref.

See also​