TeamPostgreSQL - PostgreSQL Web Admin GUI Tools
boss@solai:~$ sudo apt-get update
boss@solai:~$ sudo apt-get install libc6:i386 libncurses5:i386 libstdc++6:i386
now 32 bit lib installed in your system. try install..
A few days ago, a developer came to me with that inevitable scenario that every DBA secretly dreads: a need for a dynamic table structure. After I’d finished dying inside, I explained the various architectures that could give him what he needed, and then I excused myself to another room so I could weep silently without disturbing my coworkers. But was it really that bad? Databases have come a long way since the Bad Old Days when there were really only two viable approaches to table polymorphism. Postgres in particular adds two options that greatly reduce the inherent horror of The Blob. In fact, I might even say its a viable strategy now that Postgres JSON support is so good.
Why might a dev team need or even want a table with semi-defined columns? That seems like a silly question to ask in the days of NoSQL databases. Yet an application traditionally tied to a RDBMS can easily reach a situation where a strict table is outright inadequate. Our particular case arose because application variance demands columnar representation of precalculated values. But the generators of those values are themselves dynamic, as are the labels. One version may have ten defined columns beyond the standard core, where another might have fifteen. So we need a dynamic design of some description.
Before we go crazy in our history lesson, let’s start with a base table we’ll be expanding in various ways through the article. It’ll have a million data points ranging from today to a few months in the past, with datapoints every ten seconds.
CREATE TABLE sensor_log ( id INT PRIMARY KEY, location VARCHAR NOT NULL, reading BIGINT NOT NULL, reading_date TIMESTAMP NOT NULL ); INSERT INTO sensor_log (id, location, reading, reading_date) SELECT s.id, s.id % 1000, s.id % 100, '2016-03-11'::TIMESTAMP - ((s.id * 10) || 's')::INTERVAL FROM generate_series(1, 1000000) s(id); CREATE INDEX idx_sensor_log_location ON sensor_log (location); CREATE INDEX idx_sensor_log_date ON sensor_log (reading_date); ANAL |
After giving my presentation at ConFoo this year, I had some discussions with a few people about the ability to put constraints on JSON data, and whether any of the advanced PostgreSQL constraints work for that. Or in short, can we get the benefits from both SQL and NoSQL at the same time?
My general response to questions like this when it comes to PostgreSQL is "if you think there's a chance it works, it probably does", and it turns out that applies in this case as well.
For things like UNIQUE keys and CHECK constraints it's fairly trivial, but there are also things like EXCLUSION constraints where there are some special constructs that need to be considered.
Other than the technical side of things, it's of course also a question of "should we do this". The more constraints that are added to the JSON data, the less "schemaless" it is. On the other hand, other databases that have schemaless/dynamic schema as their main selling points, but still require per-key indexes and constraints (unlike PostgreSQL where JSONB is actually schemaless even when indexed).
Anyway, back on topic. Keys and constraints on JSON data.
In PostgreSQL, keys and constraints can be defined on both regular columns and directly on any expression, as long as that expression is immutable (meaning that the output is only ever dependent on the input, and not on any outside state). And this functionality works very well with JSONB as well.
So let's start with a standard JSONB table:
postgres=# CREATE TABLE jsontable (j jsonb NOT NULL);
CREATE TABLE
postgres=# CREATE INDEX j_idx ON jsontable USING gin(j jsonb_path_ops);
CREATE INDEX
Of course, declaring a table like this is very seldom a good idea in reality - a single table with just a JSONB field. You probably know more about your data than that, so there will be other fields in the table than just the JSONB field. But this table will suffice for our example.
A standard gin index using jsonb_path_ops is how we get fully schemaless indexing in jsonb with maximum performance. We're not actuall
[...]Good news has come out to ensure your disaster recovery strategy is even better!
Wait…what?! You don’t have a disaster recovery
strategy?! No good, my friend, no good…
But don’t despair, as I was saying, the good news is that Barman, the powerful backup and recovery manager for PostgreSQL, has just released its 1.6.0 version, which comes with important new features and, as expected, bug fixes.

So, if you are already using Barman, just update it to start using the new features. And, if you don’t yet use it, now is the time you can begin elaborating your disaster recovery plan. It’s better you don’t wait…trust me on that one. :)
“And what are those new features?” you ask. Here we go:
This is the main feature of this release. Now Barman is capable of connecting to the database server and continuously receiving WAL files via PostgreSQL’s native streaming protocol, which will reduce RPO (Recovery Point Objective). In case the streaming connection fails for any reason, the standard log archiving takes control right away, making sure WALs are being archived. This version 1.6.0 still requires that standard archiving is in place (which means you will end up transferring WAL files twice, over two different channels – but this is a small price to pay for near zero RPO).
Enabling PostgreSQL streaming connection on the Barman server is pretty straightforward. Just open the configuration file of your backup server, that is by default inside /etc/barman.d/ directory, and add the following settings:
streaming_conninfo = host=your.postgresql.server user=streaming_barman
streaming_archiver = on
archiver = on
path_prefix = /path/to/pg_receivexlog/
Each line above provides the necessary set that Barman needs to successfully use the streaming connection:
streaming_conninfo in
the same way as the already present conninfo
does:host indicates your PostgreSQL server.user informs which user will be used for this
connection. The user streaming_barman has to be
created Leo and I attended the Paris OSGeo Code Sprint at Mozilla Foundation put together by Oslandia and funded by several companies. It was a great event. There was quite a bit of PostGIS related hacking that happened by many new faces. We have detailed at BostonGIS: OSGeo Code Sprint 2016 highlights some of the more specific PostGIS hacking highlights. Giuseppe Broccolo of 2nd Quadrant already mentioned BRIN for PostGIS: my story at the Code Sprint 2016 in Paris.
At the best of times, I find it hard to generate a lot of sympathy for my work-from-home lifestyle as an international coder-of-mystery. However, the last few weeks have been especially difficult, as I try to explain my week-long business trip to Paris, France to participate in an annual OSGeo Code Sprint.
Yes, really, I “had” to go to Paris for my work. Please, stop sobbing. Oh, that was light jealous retching? Sorry about that.
Anyhow, my (lovely, wonderful, superterrific) employer, CartoDB was an event sponsor, and sent me and my co-worker Paul Norman to the event, which we attended with about 40 other hackers on PDAL, GDAL, PostGIS, MapServer, QGIS, Proj4, PgPointCloud etc.
Paul Norman got set up to do PostGIS development and crunched through a number couple feature enhancements. The feature enhancement ideas were courtesy of Remi Cura, who brought in some great power-user ideas for making the functions more useful. As developers, it is frequently hard to distinguish between features that are interesting to us and features that are using to others so having feedback from folks like Remi is invaluable.
The Oslandia team was there in
force, naturally, as they were the organizers. Because they work a
lot in the 3D/CGAL space, they
were interested in making CGAL faster, which meant they were
interested in some “expanded object header” experiments I did last
month. Basically the
EOH code allows you to return an unserialized reference to a
geometry on return from a function, instead of a flat serialiation,
so that calls that look like ST_Function(ST_Function(ST_Function()))
don’t end up with a chain of three serialize/deserialize steps in
them. When the deserialize step is expensive (as it in for their 3D
objects) the benefit of this approach is actually measureable. For
most other cases it’s not.
(The exception is in things like mutators, called from within PL/PgSQL, for example doing array appends or insertions in a tight loop. Tom Lane wrote up this enhancement of PgSQL with examples for array manipulation an
[...]In a heterogeneous database environment, it’s not uncommon for object creation and modification to occur haphazardly. Unless permissions are locked down to prevent it, users and applications will create tables, modify views, or otherwise invoke DDL without the DBA’s knowledge. Or perhaps permissions are exceptionally draconian, yet they’ve been circumvented or a superuser account has gone rogue. Maybe we just need to audit database modifications to fulfill oversight obligations. Whatever the reason, Postgres has it covered with event triggers.
Now, event triggers have only been around a relatively short
while, having appeared in version 9.3. Even though I was personally
excited to see the feature, priorities changed and they fell off my
RADAR for quite a while. Yet the functionality they offer is
exceedingly useful and worthy of exploration. So here’s a simple
scenario: email DDL (Database Definition Language) to a DBA team
whenever it occurs. This would mean anything from CREATE
TABLE to DROP RULE, or anything else
from this chart.
This would normally be fairly easy, but our first choice of
language, the native PL/pgSQL doesn’t have email
functionality. Can we use Python? Let’s try:
CREATE OR REPLACE FUNCTION sp_email_command() RETURNS event_trigger AS $$ print 'Event Trigger!' $$ LANGUAGE plpythonu; ERROR: PL/Python functions cannot RETURN TYPE event_trigger |
Nope. While this is a somewhat unfortunate oversight, it’s not a
roadblock. We can use a standard PL/pgSQL function as
a wrapper to call our Python email routine. So let’s just do
that.
The next thing to consider is what we should include in the email. Thankfully there’s a wealth of information available regarding database sessions. It’s always a good idea to be specific while snitching, so our email should minimally report who the user is, where they came from, where they went, and what they did. Here’s a very quick and dirty event trigger that does all of that:
CREATE OR REPLACE FUNCTION sp_tattle_ddl() RETURNS event_trigger AS $$ |

Partitioning in PostgreSQL is traditionally implemented using table inheritance. Table inheritance allow planner to include into plan only those child tables (partitions) which are compatible with query. Simultaneously a lot of work on partitions management remains on users: create inherited tables, writing trigger which selects appropriate partition for row inserting etc. In order to automate this work pg_partman extension was written. Also, there is upcoming work on declarative partitioning by Amit Langote for PostgreSQL core.
In Postgres Professional we notice performance problem of inheritance based partitioning. The problem is that planner selects children tables compatible with query by linear scan. Thus, for query which selects just one row from one partition it would be much slower to plan than to execute. This fact discourage many user and this is why we’re working on new PostgreSQL extension: pg_pathman.
pg_pathman caches partitions meta-information and uses set_rel_pathlist hook in order to replace mechanism of child tables selection by its own mechanism. Thanks to this binary search algorithm over sorted array is used for range partitioning and hash table lookup for hash partitioning. Therefore, time spent to partitions selection appears to be negligible in comparison with forming of result plan nodes. See postgrespro blog post for performance benchmarks.
pg_pathman now in beta-release status and we encourage all interested users to try it and give us a feedback. pg_pathman is compatible with PostgreSQL 9.5 and distributed under PostgreSQL license. In the future we’re planning to enhance functionality of pg_pathman by following features.
Despite we have pg_p
[...]Announcing: My new book, “Practice Makes Regexp,” with 50 exercises meant to help you learn and master regular expressions. With explanations and code in Python, Ruby, JavaScript, and PostgreSQL.
I spend most of my time nowadays going to high-tech companies and training programmers in new languages and techniques. Actually, many of the things I teach them aren’t really new; rather, they’re new to the participants in my training. Python has been around for 25 years, but for my students, it’s new, and even a bit exciting.
I suggested the first solution that came to mind: Regular expressions.
Regular expressions are a lifesaver for anyone who works with text. We can use them to search for patterns in files, in network data, and in databases. We can use them to search and replace. To handle protocols that have changed ever so slightly from version to version. To handle human input, which is always messier than what we get from other computers.
Regular expressions are one of the most critical tools I have in my programming toolbox. I use them at least a few times each day, and sometimes even dozens of times in a given day.
So, why don’t all developers know and use regular expressions? Quite simply, because the learning curve is so steep. Regexps, as they’re also known, are terse and cryptic. Changing one character can have a profound impact on what text a regexp matches, as well as its performance. Knowing which character to insert where,
When: 6-8pm Thursday Mar 17, 2016
Where: Iovation
Who: Jsoh Berkus
What: Big Data on Pg 9.5
PostgreSQL 9.5, released in January, is full of tasty goodness for big data folks: UPSERT, new aggregation types, OLAP support, new Foreign Data Wrapper functionality, BRIN indexes, and more. Josh Berkus will give us a rundown of these features, with multiple demos.
Josh Berkus is a member of the PostgreSQL Core Team, and works at Red Hat where he manages the Project Atomic community. Containers, containers, containers! Which is appropriate, since he’s also a potter.
—
If you have a job posting or event you would like me to announce at
the meeting, please send it along. The deadline for inclusion is
5pm the day before the meeting.
—
Our meeting will be held at Iovation, on the 32nd floor of the US Bancorp Tower at 111 SW 5th (5th & Oak). It’s right on the Green & Yellow Max lines. Underground bike parking is available in the parking garage; outdoors all around the block in the usual spots. No bikes in the office, sorry!
Elevators open at 5:45 and building security closes access to the floor at 6:30.
When you arrive at the Iovation office, please sign in on the iPad at the reception desk.
See you there!
Adoption of document databases is growing rapidly in
response to the need for solutions that can handle large volumes of
data. These solutions are adept at handling non-relational data
models, but the distinction between database types (document,
relational, key-value, etc.) is blurring. Organizations are making
broad sweeps of data from multiple sources and storing it in single
large documents for later analysis.
In the previous posts I have described a simple database table for storing JSON values, and a way to unpack nested JSON attributes into simple database views. This time I will show how to write a very simple query (thanks to PostgreSQL 9.5) to load the JSON files
Here's a simple Python script to load the database.
This script is made for PostgreSQL 9.4 (in fact it should work for 9.5 too, but is not using a nice new 9.5 feature described below).
#!/usr/bin/env python
import os
import sys
import logging
try:
import psycopg2 as pg
import psycopg2.extras
except:
print "Install psycopg2"
exit(123)
try:
import progressbar
except:
print "Install progressbar2"
exit(123)
import json
import logging
logger = logging.getLogger()
PG_CONN_STRING = "dbname='blogpost' port='5433'"
data_dir = "data"
dbconn = pg.connect(PG_CONN_STRING)
logger.info("Loading data from '{}'".format(data_dir))
cursor = dbconn.cursor()
counter = 0
empty_files = []
class ProgressInfo:
def __init__(self, dir):
files_no = 0
for root, dirs, files in os.walk(dir):
for file in files:
if file.endswith(".json"):
files_no += 1
self.files_no = files_no
print "Found {} files to process".format(self.files_no)
self.bar = progressbar.ProgressBar(maxval=self.files_no,
widgets=[' [', progressbar.Timer(), '] [', progressbar.ETA(), '] ', progressbar.Bar(),])
def update(self, counter):
self.bar.update(counter)
pi = ProgressInfo(os.path.expanduser(data_dir))
for root, dirs, files in os.walk(os.path.expanduser(data_dir)):
for f in files:
fname = os.path.join(root, f)
if not fname.endswith(".json"):
continue
with open(fname) as js:
data = js.read()
if not data:
empty_files.append(fname)
continue
import json
dd = json.loads(data)
counter += 1
[...]After months of writing, editing, and procrastinating, my new ebook, “Practice Makes Regexp” is almost ready. The book (similar to my earlier ebook, “Practice Makes Python“) contains 50 exercises to improve your fluency with regular expressions (“regexps”), with solutions in Python, Ruby, JavaScript, and PostgreSQL.
When I tell people this, they often say, “PostgreSQL? Really?!?” Many are surprised to hear that PostgreSQL supports regexps at all. Others, once they take a look, are surprised by how powerful the engine is. And even more are surprised by the variety of ways in which they can use regexps from within PostgreSQL.
I’m thus presenting an excerpt from the book, providing an overview of how PostgreSQL’s regexp operators and functions. I’ve used these many times over the years, and it’s quite possible that you’ll also find them to be of assistance when writing queries.
PostgreSQL isn’t a language per se, but rather a relational database system. That said, PostgreSQL includes a powerful regexp engine. It can be used to test which rows match certain criteria, but it can also be used to retrieve selected text from columns inside of a table. Regexps in PostgreSQL are a hidden gem, one which many people don’t even know exists, but which can be extremely useful.
Regexps in PostgreSQL are defined using strings. Thus, you will create a string (using single quotes only; you should never use double quotes in PostgreSQL), and then match that to another string. If there is a match, PostgreSQL returns “true.”
PostgreSQL’s regexp syntax is similar to that of Python and Ruby, in that you use backslashes to neutralize metacharacters. Thus, + is a metacharacter in PostgreSQL, whereas \+ is a plain “plus” character. However, there are differences between the regexp syntax —for example, PostgreSQL’s word-boundary metacharacter is \y whereas in Python and Ruby, it is \b. (This was likely done to avoid conflicts with the ASCII backspace character.)
Where things are truly different in Postgr
[...]In the previous post I showed a simple PostgreSQL table for storing JSON data. Let's talk about making the JSON data easier to use.
One of the requirements was to store the JSON from the files
unchanged. However using the JSON operators for deep attributes is
a little bit unpleasant. In the example JSON there is attribute
country inside metadata. To access this
field, we need to write:
SELECT data->'metadata'->>'country' FROM stats_data;
The native SQL version would rather look like:
SELECT country FROM stats;
So let's do something to be able to write the queries like this. We need to repack the data to have the nice SQL types, and hide all the nested JSON operators.
I've made a simple view for this:
CREATE VIEW stats AS
SELECT
id as id,
created_at as created_at,
to_timestamp((data->>'start_ts')::double precision) as start_ts,
to_timestamp((data->>'end_ts')::double precision) as end_ts,
tstzrange(
to_timestamp((data->>'start_ts')::double precision),
to_timestamp((data->>'end_ts')::double precision)
) as ts_range,
( SELECT array_agg(x)::INTEGER[]
FROM jsonb_array_elements_text(data->'resets') x) as resets,
(data->'sessions') as sessions,
(data->'metadata'->>'country') as country,
(data->'metadata'->>'installation') as installation,
(data->>'status') as status
FROM stats_data;
This is a normal view, which means that it is only a query stored in the database. Each time the view is queried, the data must be taken from the stats_data table.
There is some code I could extract to separate functions. This will be useful in the future, and the view sql should be cleaner.
Here are my new functions:
CREATE OR REPLACE FUNCTION to_array(j jsonb) RETURNS integer[] A[...]
Back in 2012 I wrote an overview of database sharding. Since
then I’ve had a few questions about it, which have really increased
in frequency over the last two months. As a result I thought I’d do
a deeper dive with some actual hands on for sharding. Though for
this hands on, because I do value my time I’m going to take
advantage of pg_shard rather than creating mechanisms
from scratch.
For those unfamiliar pg_shard is an open source extension from Citus data who has a commerical product that you can think of is pg_shard++ (and probably much more). Pg_shard adds a little extra to let data automatically distribute to other Postgres tables (logical shards) and Postgres databases/instances (physical shards) thus letting you outgrow a single Postgres node pretty simply.
Alright, enough talk about it, let’s get things up and running.
The rest assume you have Postgres.app, version 9.5 setup and are on a Mac, much of these steps could be easily adapted for other Postgres installs or OSes.
PATH=/Applications/Postgres.app/Contents/Versions/latest/bin/:$PATH make
sudo PATH=/Applications/Postgres.app/Contents/Versions/latest/bin/:$PATH make install
cp /Applications/Postgres.app/Contents/Versions/9.5/share/postgresql/postgresql.conf.sample /Applications/Postgres.app/Contents/Versions/9.5/share/postgresql/postgresql.conf.sample
Edit your postgresql.conf:
#shared_preload_libraries = ''
TO:
shared_preload_libraries = 'pg_shard'
Then create a file in /Users/craig/Library/Application\
Support/Postgres/var-9.5/pg_worker_list.conf where
craig is your username:
# hostname port-number
localhost 5432
localhost 5433
You’ll also need to create a new Postgres instance:
initdb -D /Users/craig/Library/Application\ Support/Postgres/var-9.5-2
Then edit that postgresql.conf inside that newly
created folder with two main edits:
port = 5432
To
port = 5433
Finally setup our database then start it up:
createdb instagram
postgres -D /Users/craig/Library/Application\ Support/Postgres/var-9.5-2
Now you sh
[...]
Ah, users. They log in, query things, accidentally delete critical data, and drop tables for giggles. Bane or boon, user accounts are a necessary evil in all databases for obvious reasons. And what would any stash of data be if nobody had access? Someone needs to own the objects, at the very least. So how can we be responsible with access vectors while hosting a Postgres database? We already covered automating grants, so let’s progress to the next step: building a “best practice” access stack.
More experienced (and smarter) people than me have given this
process a lot of thought, so why not learn from one of their
implementations? In UNIX systems, file access is handled through up
to nine grants based on the owner, an arbitrary group, or the
unwashed masses, each with read, write, or execute permissions. In
the context of Postgres, the shambling hordes can be associated
with PUBLIC, USER with ownership, and
GROUP for the usual buckets. Similarly, we have
SELECT for read, INSERT,
UPDATE, or DELETE for write, and
EXECUTE as well.
This means we actually have a bit more control in certain areas than UNIX filesystems. In fact, since multiple users or groups have distinct privileges on each object, we have the opportunity to be really creative. If we start with the analogous case, but expand it by leveraging multiple groups, we have a set of standardized roles that resemble a filesystem’s privileges, but with greater granularity. We can even create something similar to a bitmask, so all new objects adhere to our desired grant structure.
So how do we start? The easy case would have the owner, a group for writing, a group for reading, and a group for function execution. Let’s set it up and include the default ACLs from the previous article:
CREATE USER myapp_owner WITH PASSWORD 'test'; CREATE SCHEMA myapp AUTHORIZATION myapp_owner; ALTER USER myapp_owner SET search_path = myapp; CREATE GROUP myapp_reader; CREATE GROUP myapp_writer; CREATE GROUP myapp_exec; GRANT USAGE ON SCHEMA myapp TO myapp_reader; GRANT USAG |
A growing number of organizations are using containers, such as docker, to deploy applications and parts of their infrastructure. How well do contains work with a database such as PostgreSQL? What do we need to know about installing, configuring, and deploying PostgreSQL in this way, and what mistakes should we aim to avoid? In this talk, Jignesh Shah shares his experiences combining PostgreSQL with Docker. He describes the reasons why it’s useful to work in this way, and how we can deploy and then monitor our PostgreSQL instances in a number of ways.
Time: 1 hour, 13 minutes
The post [Video 452] Jignesh Shah: PostgreSQL and Linux Containers appeared first on Daily Tech Video.
libpq-dev is a set of library functions that allow client programs (for our case its 'R') to pass queries to the PostgreSQL backend server and to receive the results of these queries.
Earlier this week I posted an email to the hackers list outlining a plan for built-in sharding. Later emails in the thread clarify many of my initial remarks. (I also posted a presentation about this a year ago.)
The emails detail the steps being taken to enhance existing Postgres features (foreign data wrappers (FDW), parallelism, partitioning) to minimize the code changes necessary to add sharding. While there is still uncertainty about the final design, what workloads it will support, and even the probability of success, steady progress is being made. As feature enhancements are completed, the path to built-in Postgres sharding will become clearer.
The 3rd annual Nordic PGDay will be held on 17th of March in Helsinki this year. Registration is still open, reserve yourselves a seat before it’s too late!
Nordic PGDay is visiting a different Nordic city each year as a tradition. This year the conference will be in Helsinki, last year it was in Copenhagen and it was a great event! You can check last year’s Nordic PGDay to see what you’ve missed and you may not want to miss this year’s chance :)
If you are a PostgreSQL person who lives close to Nordic region or want to visit Helsinki, you should definitely check out the conference schedule. This call from the website proves that you can find something interesting for you:
Nordic PGDay 2016 is an excellent chance to learn more about the worlds most advanced open source database among your peers in the Nordic region. Whether you use PostgreSQL professionally or play with it in your spare time; whether you love the new JSONB capabilities or hack on query planner internals; Nordic PGDay will have something for you.
There will be amazing talks with variety of Postgres-related topics from the perspective of developers, DBAs, system admins, database developers and PostgreSQL users. Let’s check the talks quickly and see what will be going around this year in Helsinki.
Event will start with the registration at 08:30. There will be 7 talks, 2 coffee-breaks and a lunch break. The last talk will finish at 17:35.
Let’s see the talks and the topics.
Vik Fearing from 2ndQuadrant will give talk about “How Did We Live Without LATERAL?” between 15:45 – 16:35, since he is my colleague I will give the priority to Vik :)
In his talk he will take a simple problem and show 5 ways to write a query for it, the last one being the simplest of all using LATERAL. So if you’re using PostgreSQL in your production environment or just for fun, you would possibly want to learn tricks for optimising your queries in an efficient way. Then this talk might be interestin
[...]We have plenty of Liquid Galaxy systems, where we write statistical information in json files. This is quite a nice solution. However we end with a bunch of files on a bunch of machines.
Inside we have a structure like:
{
"end_ts": 1438630833,
"resets": [],
"metadata": {
"country": "USA",
"installation": "FIRST"
},
"sessions": [
{
"application": "first",
"end_ts": 1438629089,
"start_ts": 1438629058
},
{
"application": "second",
"end_ts": 1438629143,
"start_ts": 1438629123
},
{
"application": "third",
"end_ts": 1438629476,
"start_ts": 1438629236
}
],
"start_ts": 1438629033,
"status": "on"
}
And the files are named like "{start_ts}.json". The number of files is different on each system. For January we had from 11k to 17k files.
The fields in the json mean:
We keep these files in order to get statistics from them. So we can do one of two things: keep the files on disk, and write a script for making reports. Or load the files into a database, and make the reports from the database.
The first solution looks quite simple. However for a year of files, and a hundred of systems, there will be about 18M files.
The second solution has one huge advantage: it should be faster. A database should be able to have some indexes, where the precomputed data should be stored for faster querying.
For a database we chose PostgreSQL. The 9.5 version released in January has plenty of great features for managing JSON data.
The basic idea behind the database schema is:
JD Wrote:
Once again United States PostgreSQL is present at Linux Fest Northwest. We are celebrating Free and Open Source software with 2000 other advocates in what is the second largest Free & Open Source conference on the West cost. We have a great list of talks this year with new speakers, speakers from both SeaPUG, WhatcomPUG and PDXPUG. Will you be joining us? Here is the list of PostgreSQL talks.