Hacker Newsnew | past | comments | ask | show | jobs | submit | keshavmr's commentslogin

"Six thousand years ago, Sumarians invented writing for transaction processing."... is the first sentence of the "Transaction Processing" book... https://www.amazon.com/Transaction-Processing-Concepts-Techn...


Wow that is the book to read for legacy systems


No, it is proposing an ambitious rearchitecture of current systems on new principles which it argues were present in an incomplete form in many existing systems. Since it was written, some of its proposals have been implemented to generally good effect, but by no means all. It's a visionary manifesto for the future, though filled with practical information.


At the 2012 Turning Award conference in San Francisco, Prof William Kahan mentioned that he had a newer test suite available in 1993 that would have caught Intel's bug. Still, Intel did not run that.. Prof. Kahan was actively involved in its analysis and further testing. (I'm stating this just from memory).


In this article, we’ll give you an overview of the challenges of implementing column-wise storage for JSON and the techniques used to address these challenges.


Capella Columnar is an advanced real-time analytics database service from Couchbase, targeted for real-time data processing, offering SQL++ for processing JSON (semi-structured) data and more. This service enables data to be managed locally and streamed continuously from both relational and NoSQL databases, or simply process data on S3. The columns or fields of the source are directly mapped to a field in the JSON document at the destination automatically. This is really a zero-ETL operation. A key feature of this system is its ability to continuously stream data, making it immediately available for querying, thus ensuring near real-time data processing. The JSON analytics database engine is designed for MPP (Massively Parallel Processing), with a column store for JSON and a cost-based optimizer.


bleve is a FTS engine written fully in go. https://blevesearch.com/


And, going for the HN gold, Toshi is a FTS server written in Rust: https://github.com/toshi-search/Toshi#readme (backed by the FTS library https://github.com/tantivy-search/tantivy#readme also in Rust)


Adding to this, here's a couple more options (also written in Rust):

Meilisearch - https://meilisearch.com/ Sonic - https://github.com/valeriansaliou/sonic#readme

Sonic is slightly different in that it returns the indexes to the documents that you use to go and get from an external store, but the engineering behind it is super cool: it uses finite-state-transducers to build the trie's which are then search.

Edit: updated to correct link


Thanks for reminding me about MeiliSearch!

I think you meant your 2nd link to be https://github.com/valeriansaliou/sonic#readme


Yes I certainly did, thanks for the catch.


SELECT vendor_id, cab_type, avg(fare_amount) from trips;

This takes ~86 seconds. Ran it multiple times.

SIMD is one of the ingredients for better query performance, but NOT THE ONLY ONE. See this for more info: https://www.youtube.com/watch?v=xJd8M-fbMI0


Correct. Couchbase uses Bleve to provide the distributed search index/query. See: https://docs.couchbase.com/server/6.0/fts/full-text-intro.ht...


Given a schema, you can generate an infinite number of queries. The starting from the business question is a better choice. There are some interesting work done in this area, but still a long way to go. Here's an example: https://einstein.ai/static/images/pages/research/seq2sql/seq...


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: