# Import articles in bulk

_Category: Guides and recipes_

Move articles from another system into HelpCenter.io: put them in a CSV or JSON file, then run the Python script on this page. It maps your categories, sends the articles 50 at a time, copies their images into your help center, and can run again without creating duplicates.

## Before you start

- You need a **Read & write** API key for the help center. Owners and Admins can create one: see [Create an API key](https://self.helpcenter.io/content/create-an-api-key). The API works on every plan.
- You need Python 3.8 or later. The script uses the standard library only.
- Every language in your file must be enabled in the help center before you import: see [Add a language to your help center](https://self.helpcenter.io/content/add-a-language).
- Moving a whole help center from another platform? Our migration team can do it for you: see [Can you move my existing help center to HelpCenter.io?](https://self.helpcenter.io/content/can-you-migrate-my-existing-help-center-to-helpcenter.io)

## Step 1: Put your articles in a file

Export your articles from the old system, then shape them into a CSV file with a header row, or a JSON array of objects. Each row or object is one article in one language:

| Field | What it does |
| --- | --- |
| `id` | Required. The article's id in the old system. The script sends it as the article's `external_id`, `import:<id>`, so a second run updates the article instead of adding a copy. |
| `title` | Required. |
| `body` | The article, as HTML, or as Markdown when `format` is `markdown`. |
| `category` | A category name, or a path such as `Billing > Invoices` for a subcategory. Categories that don't exist yet are created. Without it, a new article is filed under Uncategorized. |
| `published` | `true` or `false`. Leave it empty to create the article as a draft, and to leave its published state alone on later runs. |
| `format` | `html`, the default, or `markdown`. |
| `language` | A language code, such as `de`. Rows with the same `id` become one article in several languages. Empty means your help center's default language. |

This `articles.csv` holds three articles: one in English and German, one written in Markdown with an image on another server, and one draft:

```
id,language,title,body,category,published,format
101,en,Reset your password,"<p>Open <strong>Settings</strong>, then <strong>Security</strong>, and click <strong>Reset password</strong>.</p>",Account,true,html
101,de,Passwort zurücksetzen,"<p>Öffnen Sie <strong>Einstellungen</strong>, dann <strong>Sicherheit</strong>, und klicken Sie auf <strong>Passwort zurücksetzen</strong>.</p>",Account,true,html
102,en,Format invoice notes with Markdown,"Invoice notes support **bold**, *italic* and lists.

![The Markdown mark](https://raw.githubusercontent.com/github/explore/main/topics/markdown/markdown.png)",Billing > Invoices,true,markdown
103,en,Change your plan,<p>Open <strong>Billing</strong> and click <strong>Change plan</strong>.</p>,Billing,,html
```

## Step 2: Save the script

Save this as `import_articles.py`:

```
#!/usr/bin/env python3
"""Import articles into a HelpCenter.io help center from a CSV or JSON file.

    HELPCENTER_API_KEY=... python3 import_articles.py articles.csv

One row (CSV) or object (JSON) per article and language, with these fields:
  id         required: the article's id in the old system (it keys re-runs)
  title      required
  body       the article as HTML, or as Markdown when format is "markdown"
  category   optional: "Billing", or "Billing > Invoices" for a subcategory
  published  optional: true or false (a new article without it is a draft)
  format     optional: html (the default) or markdown
  language   optional: a language code; rows with the same id become one article
Python 3.8 or later, standard library only.
"""
import csv
import json
import os
import sys
import time
import urllib.error
import urllib.request

API = 'https://api.helpcenter.io/v1'
KEY = os.environ.get('HELPCENTER_API_KEY', '')
ID_PREFIX = 'import:'  # external_id "import:<id>": a re-run updates, not duplicates
BATCH_SIZE = 50  # the most POST /v1/articles/bulk accepts
MAX_BATCH_BYTES = 8_000_000  # stay well under the 32 MB request limit
BULK_INTERVAL = 3  # seconds between bulk calls: at most 20 a minute from one IP
RETRY_CODES = {'external_id_busy', 'persist_failed'}

def api(method, path, body=None):
    """Send one request and return (status, body). Waits and retries on 429."""
    data = None if body is None else json.dumps(body).encode()
    headers = {'apikey': KEY, 'Accept': 'application/json'}
    if data is not None:
        headers['Content-Type'] = 'application/json'
    for attempt in range(1, 9):
        request = urllib.request.Request(
            API + path, data=data, method=method, headers=headers)
        try:
            with urllib.request.urlopen(request, timeout=300) as response:
                return response.status, json.load(response)
        except urllib.error.HTTPError as error:
            retry_after = error.headers.get('Retry-After', '')
            if error.code == 429 and attempt < 8:
                wait = int(retry_after) if retry_after.isdigit() else 2 ** attempt
                print(f'Rate limited on {method} {path}, retrying in {wait} s')
                time.sleep(wait)
                continue
            text = error.read().decode('utf-8', 'replace')
            try:
                return error.code, json.loads(text)
            except ValueError:
                return error.code, {'message': text[:300]}

def expect(result, statuses, what):
    status, body = result
    if status not in statuses:
        sys.exit(f'{what}: HTTP {status} {json.dumps(body)[:300]}')
    return body

def read_rows(path):
    if path.endswith('.json'):
        with open(path, encoding='utf-8') as file:
            data = json.load(file)
        return data['articles'] if isinstance(data, dict) else data
    with open(path, newline='', encoding='utf-8-sig') as file:
        return list(csv.DictReader(file))

def value(row, field):
    raw = row.get(field)
    return '' if raw is None else str(raw).strip()

def category_resolver(language):
    """Map a name or "Parent > Child" path to a category id, creating missing ones."""
    categories, page = [], 1
    while True:
        body = expect(api('GET', f'/categories?limit=100&page={page}'), [200],
                      'List categories')
        categories += body['categories']
        if page >= body['meta']['total_pages']:
            break
        page += 1

    def named(category):
        return ((category['name'] or {}).get(language) or '').strip().lower()

    def category_id(path):
        parent = None
        for name in [part.strip() for part in path.split('>') if part.strip()]:
            match = next((c for c in categories
                          if c['parent'] == parent and named(c) == name.lower()), None)
            if match is None:
                payload = {'name': {language: name}}
                if parent:
                    payload['parent_id'] = parent
                created = expect(api('POST', '/categories', payload), [201],
                                 f'Create category "{name}"')
                match = dict(created['category'], parent=parent)
                categories.append(match)
                print(f'Created category "{name}"')
            parent = match['id']
        return parent

    return category_id

def build_articles(rows, default_language, category_id):
    articles = {}
    for number, row in enumerate(rows, start=1):
        source_id = value(row, 'id')
        if not source_id:
            sys.exit(f'Row {number} has no id.')
        language = value(row, 'language') or default_language
        article = articles.setdefault(source_id, {
            'external_id': ID_PREFIX + source_id,
            'rehost_images': True,  # copy images from the old system to HelpCenter.io
        })
        # Empty cells are left out, so the API reports a missing title by name.
        for field, column in (('title', 'title'), ('content', 'body')):
            if value(row, column):
                article.setdefault(field, {})[language] = value(row, column)
        if value(row, 'format').lower() == 'markdown':
            article['content_format'] = 'markdown'
        if value(row, 'category') and 'category_id' not in article:
            article['category_id'] = category_id(value(row, 'category'))
        published = value(row, 'published').lower()
        if published:
            article['published'] = published in ('true', '1', 'yes')
    return list(articles.values())

def batches(articles):
    batch, size = [], 0
    for article in articles:
        article_size = len(json.dumps(article))
        full = len(batch) == BATCH_SIZE or size + article_size > MAX_BATCH_BYTES
        if batch and full:
            yield batch
            batch, size = [], 0
        batch.append(article)
        size += article_size
    if batch:
        yield batch

def send(articles, counts):
    """Send articles in bulk batches; return the ones that failed, with their errors."""
    failed, last_call = [], 0.0
    for batch in batches(articles):
        time.sleep(max(0.0, last_call + BULK_INTERVAL - time.monotonic()))
        last_call = time.monotonic()
        body = expect(api('POST', '/articles/bulk', {'articles': batch}), [207],
                      'Bulk import')
        for result in body['results']:
            article = batch[result['index']]
            if result['status'] == 'error':
                failed.append((article, result['error']))
                continue
            counts[result['status']] += 1
            images = result.get('images', {})
            counts['copied'] += images.get('rehosted', 0)
            counts['reused'] += images.get('reused', 0)
            for image in images.get('failed', []):
                print(f"{article['external_id']}: image not copied "
                      f"({image['reason']}): {image['src']}")
    return failed

def main():
    if len(sys.argv) != 2:
        sys.exit('Usage: python3 import_articles.py <file.csv or file.json>')
    if not KEY:
        sys.exit('Set HELPCENTER_API_KEY to a Read & write API key.')
    rows = read_rows(sys.argv[1])
    sites = expect(api('GET', '/sites'), [200], 'Read the help center')['sites']
    language = sites[0]['default_language']
    articles = build_articles(rows, language, category_resolver(language))

    counts = {'created': 0, 'updated': 0, 'copied': 0, 'reused': 0}
    failed = send(articles, counts)
    retry = [article for article, error in failed if error['code'] in RETRY_CODES]
    if retry:
        print(f'Retrying {len(retry)} article(s) in 10 s')
        time.sleep(10)
        failed = [item for item in failed if item[1]['code'] not in RETRY_CODES]
        failed += send(retry, counts)

    for article, error in failed:
        details = '; '.join(f"{field}: {' '.join(messages)}"
                            for field, messages in error.get('errors', {}).items())
        print(f"Failed {article['external_id']}: {error['code']} "
              f"{details or error['message']}")
    print(f"{len(articles)} articles: {counts['created']} created, "
          f"{counts['updated']} updated, {len(failed)} failed. "
          f"Images: {counts['copied']} copied, {counts['reused']} reused")
    sys.exit(1 if failed else 0)

if __name__ == '__main__':
    main()
```

## Step 3: Run it

With your key in the `HELPCENTER_API_KEY` environment variable, pass the file to import:

```
export HELPCENTER_API_KEY="paste-your-key-here"
python3 import_articles.py articles.csv
```

The first run creates the categories it can't find and then the articles:

```
Created category "Account"
Created category "Invoices"
3 articles: 3 created, 0 updated, 0 failed. Images: 1 copied, 0 reused
```

Billing already existed in this help center, so the script reused it. It created Account, and Invoices under Billing. The Markdown image now lives in your help center, so it keeps working after you close the old system.

Run the same command again, and the same three articles are updated, not copied:

```
3 articles: 0 created, 3 updated, 0 failed. Images: 0 copied, 1 reused
```

The second run downloads the image again, finds that the help center already has it, and reuses the stored copy.

## Step 4: Fix what failed, then run again

When the API refuses an article, the script names it and says why, imports the rest, and exits with status 1. In this `articles.json`, the second article has no title:

```
[
  {
    "id": 104,
    "title": "Download a receipt",
    "body": "<p>Open <strong>Billing</strong>, then <strong>Receipts</strong>.</p>",
    "category": "Billing",
    "published": true
  },
  {
    "id": 105,
    "title": "",
    "body": "<p>This one has no title.</p>",
    "category": "Billing"
  }
]
```

The run reports it:

```
Failed import:105: validation_error title: The title field is required.
2 articles: 1 created, 0 updated, 1 failed. Images: 0 copied, 0 reused
```

Fix the file and run it again: the articles that went in the first time are updated, and the fixed one is created. The script also retries, once and after 10 seconds, any article that failed with `external_id_busy` or `persist_failed`, because those can pass on a second try.

## Step 5: Review, publish and release translations

- Articles imported without `published` are drafts. Review them in the dashboard, then publish them there, or set `published` to `true` in the file and run the import again.
- Translations are stored, but readers who open an article in another language get a not-found page until you release that translation: see [Release translations and change your default language](https://self.helpcenter.io/content/manage-languages).
- Once you've switched over, stop running the import. Every run sends every row again, so it replaces edits you made in the dashboard since.

## How it works

### One request for up to 50 articles

The script sends the articles to `POST /v1/articles/bulk`, in batches of up to 50, which is the most one request can hold. Each article takes the same fields as `POST /v1/articles`: see [Articles](https://developers.helpcenter.io/content/articles-api). For example, this request sends two articles. The first has two images, and the second has no title:

```
{
  "articles": [
    {
      "external_id": "import:201",
      "rehost_images": true,
      "title": { "en": "Add a note to an invoice" },
      "content": {
        "en": "Notes support **Markdown**.\n\n![The Markdown mark](https://raw.githubusercontent.com/github/explore/main/topics/markdown/markdown.png)\n\n![The note field](https://old-help.example.com/images/note-field.png)"
      },
      "content_format": "markdown",
      "category_id": 343,
      "published": true
    },
    {
      "external_id": "import:202",
      "rehost_images": true,
      "content": { "en": "<p>This one has no title.</p>" },
      "category_id": 340
    }
  ]
}
```

The API answers `207 Multi-Status`, with one result per article, in the order you sent them. Each `article` is the full article object, shortened here:

```
{
  "status": "success",
  "results": [
    {
      "index": 0,
      "status": "created",
      "article": {
        "id": 809,
        "external_id": "import:201",
        "title": { "en": "Add a note to an invoice" },
        "published": true,
        "_links": {
          "view": {
            "method": "GET",
            "url": "https://acme.helpcenter.io/content/add-a-note-to-an-invoice"
          }
        }
      },
      "images": {
        "rehosted": 0,
        "reused": 1,
        "failed": [
          {
            "src": "https://old-help.example.com/images/note-field.png",
            "reason": "the host could not be resolved"
          }
        ],
        "failed_count": 1
      }
    },
    {
      "index": 1,
      "status": "error",
      "external_id": "import:202",
      "error": {
        "code": "validation_error",
        "message": "The article could not be accepted.",
        "errors": { "title": ["The title field is required."] }
      }
    }
  ],
  "meta": { "total": 2, "created": 1, "updated": 0, "failed": 1 }
}
```

The status is `207` even when every article failed, so read each result:

- `index` is the article's position in your request, and `status` is `created`, `updated` or `error`.
- `images` appears when `rehost_images` is on: `rehosted` counts images copied now, `reused` images the help center already had, and `failed` lists the ones that couldn't be copied, with the reason.
- `meta` counts the results.

An `error` result carries one of these codes:

| Code | What it means | What to do |
| --- | --- | --- |
| `validation_error` | A field broke a rule. `errors` lists each field with its messages. | Fix the data. |
| `invalid_argument` | A language that isn't enabled in the help center: `Unsupported language key "de" provided.` | Add the language in the dashboard, or leave it out. |
| `invalid_item` | The item isn't a JSON object. | Send an object per article. |
| `external_id_busy` | Another request is importing the same `external_id` right now. | Retry after a few seconds. |
| `persist_failed` | The article couldn't be saved. A new article sent with `categories`, or with a title that is an empty string, fails this way too. | Retry once. Leave `categories` and empty titles out. |

Some problems refuse the whole request, and nothing is imported. More than 50 articles gets `422 Unprocessable Entity`:

```
{
  "status": "validation_error",
  "message": "The request could not be accepted.",
  "errors": {
    "articles": ["Too many articles in one request. The maximum is 50 per batch."]
  }
}
```

A **Read only** key gets `403 Forbidden` with `{"status":"error","code":"insufficient_scope","message":"This action requires the content.write scope."}`, and a wrong key gets `401 Unauthorized` with `{"status":"unauthorized"}`. The script stops on these and prints the response.

### Categories

Before it sends any article, the script lists your categories with `GET /v1/categories` and matches each name in your default language, under the right parent. It creates the missing ones with `POST /v1/categories`, passing `parent_id` for a subcategory. Categories created through the API are public. To import into a private category, create it in the dashboard before you run the script, and the script uses it (see [Make a category private](https://self.helpcenter.io/content/private-categories)).

The import puts each article in one category, with `category_id`. To show an article in several categories, send `categories` in a `PATCH` after your last import run:

```
curl -X PATCH https://api.helpcenter.io/v1/articles/4711 \
  -H "apikey: $HELPCENTER_API_KEY" \
  -H "Accept: application/json" \
  -H "Content-Type: application/json" \
  -d '{"categories": [340, 342]}'
```

The article then reports `"category_id": -1` and `"categories": [340, 342]`.

### Images

With `rehost_images: true`, HelpCenter.io downloads every image in the article's HTML whose `src` is an absolute `http` or `https` address, or a `data:` URI, stores it in your help center and rewrites the address. For Markdown, that happens after the conversion to HTML. The limits:

- JPEG, PNG, GIF and WebP images of up to 30 MB.
- At most 25 images per article, 10 seconds per download, and 45 seconds for all the downloads in one request.
- Relative addresses and `srcset` aren't copied.

An image that can't be copied keeps its original address, and `images.failed` says why. When you run the import again, HelpCenter.io tries those images again, and reuses the ones it already stored.

If your export has the images as files instead of links, upload them first with `POST /v1/images` and put the returned URLs in the bodies before you import. [Sync articles from a Git repository](https://developers.helpcenter.io/content/sync-articles-from-git) does exactly that, and [Images](https://developers.helpcenter.io/content/images-api) has the details.

### Running it again

- `external_id` is what matches a row to its article. Keep `ID_PREFIX` and your ids the same between runs: change them, and the next run creates copies. When you import from two systems, give each its own prefix.
- An `external_id` can be up to 191 characters.
- Never send a title that is an empty string. A new article then fails with `persist_failed`, and an existing one loses its title in that language. The script leaves empty cells out for this reason.
- If an imported article is moved to the Trash, the next run creates a new article for its row.
- Every imported article has the person who created the API key as its author.

### Rate limits

HelpCenter.io counts bulk requests, image uploads and export pages from one IP address together, and refuses a bulk request once that count reaches 20 in a minute. The script waits 3 seconds between bulk requests, so it never sends more than 20 a minute. When the API still answers `429 Too Many Requests`, because something else on the same IP address used the count, the script waits the seconds in the `Retry-After` header and tries again.

Bulk requests don't count against your key's 120 writes a minute, but creating a category does. The script also starts a new request when a batch would pass 8 MB, well under the 32 MB the API accepts. See [Rate limits](https://developers.helpcenter.io/content/rate-limits).

## Troubleshooting

**`invalid_argument Unsupported language key "de" provided.`** The help center doesn't have that language yet. Add it (see [Add a language to your help center](https://self.helpcenter.io/content/add-a-language)), then run the import again.

**A translated article shows a not-found page.** The translation hasn't been released. See [Release translations and change your default language](https://self.helpcenter.io/content/manage-languages).

**`image not copied (the host could not be resolved)`** The old server is gone or unreachable. Upload the image yourself and update the article, or fix the address in the file and run again.

**`image not copied (the request ran out of time to re-host more images)`** One request holds more image downloads than 45 seconds allow. Lower `BATCH_SIZE`, to 10 for example, and run the import again.

**`Rate limited on POST /articles/bulk`** Something else sent bulk requests, image uploads or export requests from the same IP address in the same minute. The script waits and carries on by itself.

**`HTTP 403` with `"code":"insufficient_scope"`** The key is **Read only**. Use a **Read & write** key.

**The help center has two copies of an article.** The row's `external_id` changed between runs: a different `ID_PREFIX` or id, or the first copy was moved to the Trash. Delete the copy you don't want.

## Related

- [Articles](https://developers.helpcenter.io/content/articles-api)
- [Images](https://developers.helpcenter.io/content/images-api)
- [Categories](https://developers.helpcenter.io/content/categories-api)
- [Rate limits](https://developers.helpcenter.io/content/rate-limits)
- [Sync articles from a Git repository](https://developers.helpcenter.io/content/sync-articles-from-git)
