# Better organisation for large data sets

**URL:** <https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570>\
**Category:** Questions\
**Created:** [September 27, 2024, 9:59am UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570 "2024-09-27T09:59:12Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![breaddes](https://avatars.discourse-cdn.com/v4/letter/b/e47c2d/32.png) [@breaddes](https://forum.getkirby.com/u/breaddes)\
**Post date:** [September 27, 2024, 9:59am UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/1 "2024-09-27T09:59:12Z")

</div>

I have records of books. That’s about 1500 subpages of the books page. On the customer server, the queries run for 3-4 seconds until a result comes back from the server.

Is there a better organistation solution if I want to keep the possibility to search all 1500 records or does Kirby reach a limit here if the server does not provide the disk performance?

---

<div class="post-metadata">

**Author:** ![bastianallgeier](https://dub1.discourse-cdn.com/flex017/user_avatar/forum.getkirby.com/bastianallgeier/32/7569_2.png) [@bastianallgeier](https://forum.getkirby.com/u/bastianallgeier)\
**Post date:** [September 27, 2024, 10:13am UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/2 "2024-09-27T10:13:17Z")

</div>

If you can categorize books, that’s a first great step to get more performance and a better organisation in the panel. You could organise them by authors for example. E.g.

```auto
/books
  /author-a
    /book-a
    /book-b
  /author-b
    /book-c
    /book-d

```

For larger data sets, it is really worth looking into an external search service. Something like a self-hosted elastic search instance or a service like Algolia. We use this for [getkirby.com](http://getkirby.com) as well. You don’t just get better performance, you also get a much better search experience and powerful query options.

---

<div class="post-metadata">

**Author:** ![breaddes](https://avatars.discourse-cdn.com/v4/letter/b/e47c2d/32.png) [@breaddes](https://forum.getkirby.com/u/breaddes)\
**Post date:** [September 27, 2024, 1:07pm UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/3 "2024-09-27T13:07:01Z")

</div>

I can categorise them by genre. Does this mean that searching through several smaller directories is faster than through one large directory?

---

<div class="post-metadata">

**Author:** ![bastianallgeier](https://dub1.discourse-cdn.com/flex017/user_avatar/forum.getkirby.com/bastianallgeier/32/7569_2.png) [@bastianallgeier](https://forum.getkirby.com/u/bastianallgeier)\
**Post date:** [September 27, 2024, 3:42pm UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/4 "2024-09-27T15:42:54Z")

</div>

If you want to search through all books, the reorganisation wouldn’t make the search queries faster. But you would gain massive performance benefits when navigating through the panel to edit books.

That’s why I suggested an additional search service. You would then create a complete search index for all your books with that search service and the search would be super fast, no matter how many books you store.

---

<div class="post-metadata">

**Author:** ![carstengrimm](https://dub1.discourse-cdn.com/flex017/user_avatar/forum.getkirby.com/carstengrimm/32/188_2.png) [@carstengrimm](https://forum.getkirby.com/u/carstengrimm)\
**Post date:** [September 29, 2024, 4:11pm UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/5 "2024-09-29T16:11:29Z")

</div>

you could check out the plugins:

> **[Boost](https://plugins.getkirby.com/bnomei/boost)**
>
> Boost the speed of Kirby by having content files of pages cached and a fast lookup based on uuids

> **[Lapse](https://plugins.getkirby.com/bnomei/lapse)**
>
> Cache any data until the set expiration time

@bnomei

which might be useful for a large collection of pages, in the docs there are also a lot of good hints when dealing with large collections.

maybe you could try to cache the index of your books unless modified.

otherwise you could also try another caching driver such as apcu or sqlite.

---

<div class="post-metadata">

**Author:** ![breaddes](https://avatars.discourse-cdn.com/v4/letter/b/e47c2d/32.png) [@breaddes](https://forum.getkirby.com/u/breaddes)\
**Post date:** [October 1, 2024, 9:25am UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/6 "2024-10-01T09:25:25Z")

</div>

Thank you for your suggestions @bastianallgeier @carstengrimm. I will look into it.

Bastian in your basic content folder video you explain that blog entry directories could also be prefixed with e.g. 20241001\_.

Is it easy to implement this so that the date is used as a prefix when saving an entry? In this case, I could already filter extensive data records using this directory prefix, e.g. only books from the year 2024. I could also dispense with manual sorting here and set it completely to date.

---

<div class="post-metadata">

**Author:** ![bastianallgeier](https://dub1.discourse-cdn.com/flex017/user_avatar/forum.getkirby.com/bastianallgeier/32/7569_2.png) [@bastianallgeier](https://forum.getkirby.com/u/bastianallgeier)\
**Post date:** [October 1, 2024, 2:28pm UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/7 "2024-10-01T14:28:09Z")

</div>

Yes, you can use the num option in the books’ blueprint to set the prefix based on a date field: [Page blueprint | Kirby CMS](https://getkirby.com/docs/reference/panel/blueprints/page#sorting)

---

<div class="post-metadata">

**Author:** ![bnomei](https://dub1.discourse-cdn.com/flex017/user_avatar/forum.getkirby.com/bnomei/32/775_2.png) [@bnomei](https://forum.getkirby.com/u/bnomei)\
**Post date:** [October 1, 2024, 7:18pm UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/8 "2024-10-01T19:18:17Z")

</div>

I agree with Bastian that using an external service to speed up the search is a very efficient way to increase search performance. Alternatively, as Carsten suggested, you could speed up the Kirby instance using one of my plugins.

1. Adding a wrapping cache around your front-end search results will speed up recurring searches.

Taken from one of my projects this is how I cache search results…

**site/templates/search.php**

```php
<?php

$query = get('q', '');
$results = null;

if (strlen($query) >= 3) {
    $language = kirby()->language()->code();
    $data = lapse('search' . $query . $language, function () use ($query) {
        $results = site()->index()->listed()->search($query, explode('|', 'title|text|blocks'))->limit(7);
        return array_values($results->toArray(function ($page) {
            $field = $page->text()->or($page->blocks());
            if (Str::startsWith($field->value(), '[{"')) {
                $field = $field->toBlocks();
            }
            return [
              'title' => $page->title()->value(),
              'url' => $page->url(),
              'excerpt' => Str::unhtml(markdown($field->excerpt(200))),
            ];
        }));
    }, 15); // in minutes

    $results = new \Kirby\Cms\Collection($data);
}

if ($results && $results->count()): ?>
  <ul>
    <?php foreach ($results as $result): ?>
      <li>
        <a href="<?= $result['url'] ?>"><?= $result['title'] ?></a>
        <p><?= $result['excerpt'] ?></p>
      </li>
    <?php endforeach ?>
  </ul>
<?php endif;

```

1. The benefit of adding a plugin like _Boost_ is that every query in your front-end and within the panel, not just searches, will be quicker (3x-4x) as the content is loaded from RAM (in the case of APCu) instead of files.

---

<div class="post-metadata">

**Author:** ![breaddes](https://avatars.discourse-cdn.com/v4/letter/b/e47c2d/32.png) [@breaddes](https://forum.getkirby.com/u/breaddes)\
**Post date:** [October 2, 2024, 2:28pm UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/9 "2024-10-02T14:28:17Z")

</div>

I thought it was a bit too simple. In the blueprint, the num field in the filter doesn’t seem to be available? Is that correct?

`query: page.children.filter('num', 'num <', '20241015')`

Is there a way that I can filter pages by their sort number?

---

<div class="post-metadata">

**Author:** ![bnomei](https://dub1.discourse-cdn.com/flex017/user_avatar/forum.getkirby.com/bnomei/32/775_2.png) [@bnomei](https://forum.getkirby.com/u/bnomei)\
**Post date:** [October 2, 2024, 6:11pm UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/10 "2024-10-02T18:11:38Z")

</div>

> [@breaddes](#):
>
> query: page.children.filter(‘num’, ‘num \<’, ‘20241015’)

**filterBy and just \< as operator**  
`query: page.children.filterBy('num', '<', '20241015')`

> **[Filtering collections](https://getkirby.com/docs/cookbook/collections/filtering#filtering-using-filter-operators)**
>
> Filter pages, files and users with Kirby's extensive filtering methods.

---

<div class="post-metadata">

**Author:** ![breaddes](https://avatars.discourse-cdn.com/v4/letter/b/e47c2d/32.png) [@breaddes](https://forum.getkirby.com/u/breaddes)\
**Post date:** [October 4, 2024, 3:02pm UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/11 "2024-10-04T15:02:14Z")

</div>

That’s great! Can ‘20241015’ be set dynamically?

---

<div class="post-metadata">

**Author:** ![bnomei](https://dub1.discourse-cdn.com/flex017/user_avatar/forum.getkirby.com/bnomei/32/775_2.png) [@bnomei](https://forum.getkirby.com/u/bnomei)\
**Post date:** [October 4, 2024, 3:31pm UTC](https://forum.getkirby.com/t/better-organisation-for-large-data-sets/32570/12 "2024-10-04T15:31:22Z")

</div>

sure. any php variable containing a string or integer will work.
