Congrats, that's incredible news. Watched apavlo@'s lectures while studying in the university and finished my bachelor thesis implementing features and doing research at ClickHouse. Surprised to see these worlds being together now!
> The goal of ClickHouse Labs is to establish a best-in-class industry research organization focused on databases. It will not operate as an isolated research organization that throws ideas over the wall to engineering. Instead, we will work closely with ClickHouse engineers, customers, collaborators, and industry partners to develop and disseminate new ideas that keep ClickHouse at the bleeding edge.
This is cool, though bittersweet that the public research infrastructure (universities) is not really configured to support this kind of high-impact research any more.
Deep tech fields have a more complex relationship with academia than other industries. Engineers often need license to test ideas that may not see adoption for years, or require tens of millions to build-out.
A "research" arm is exactly this license, although it comes at the cost of potentially killing innovation in the rest of the company.
C'mon, databases are great but they are not in any way a "deep tech" field. Databases exist today in a zillion production forms as commercial products and free software.
Deep tech includes things like nuclear fusion, solid state batteries, quantum computers. I know everyone wants to feel cool, but just because your new javascript framework will be in beta for the next ten years doesn't make it "deep tech".
The databases we use today in production have severe limitations and are not even close to what is theoretically possible. Many traditional parts of a database (indexing, caching, scheduling, et al) are AI-complete algorithm problems. Entire sub-classes of database (e.g. graph or spatial) famously have persistently poor scalability and performance because of open questions in the foundational computer science.
Just the fact that increasing the generality, scalability, and performance of databases asymptotically converges on designing AGI suggests that it is, in fact, "deep tech". And this property has to mesh with other practical constraints on database behavior. Many problems in databases are hard with little forward progress in decades.
It is true that most database research is not deep tech but there is ample room for it to be if one is sufficiently ambitious.
Could you explain what practical research there is to be done? The heavy theory I know does not seem to be very useful in practice. Optimal join algorithms, Yannakakis adjacent algorithms, tree decomposition of queries all seem to be worse than well implemented naive algorithms. But maybe the implementations of the new algorithms just are not good? I really don’t know.
I'd ascribe "deep tech" to anything that you can reasonably get a PhD in and have it not be unusual. There are dozens of academic conferences on DBs pushing the frontier forward.
I'm not sure what you think "research lab" means? As I mentioned in the article, we look at IBM Research (Almaden) and Microsoft Research as inspiration.
Many database companies have research labs. It allows you to explore ideas, unorthodox concepts, and parts of the design space that may never be reflected in an actual product. Or to figure out how to solve specific hard problems that you come across with neither a good solution nor proof of impossibility in literature.
This is essential for a company that wants to stay on the frontier of database tech.
Which database companies have research labs? AFAIK amazon, snowflake, databricks, google and SAP don’t have a research lab dedicated to databases. They have some people that are paid to do research but there does not seem to be an IBM Almaden anywhere in the world.
> ClickHouse had features that at the time were only found in a handful of closed-source, commercial analytical DBMSs. For example, ClickHouse was written in C++ and supported vectorized query execution using SIMD in 2016. Most prominent open-source analytical DBMSs in 2016 were JVM-based and did not support SIMD optimizations until years later.
Performance is a feature. "Written in C++" is a strange idea of a feature.
What idea are you most excited about to work on first?
This is cool, though bittersweet that the public research infrastructure (universities) is not really configured to support this kind of high-impact research any more.
Whatever floats your boat. Sounds like you just work as an engineer at a db company
A "research" arm is exactly this license, although it comes at the cost of potentially killing innovation in the rest of the company.
Deep tech includes things like nuclear fusion, solid state batteries, quantum computers. I know everyone wants to feel cool, but just because your new javascript framework will be in beta for the next ten years doesn't make it "deep tech".
Just the fact that increasing the generality, scalability, and performance of databases asymptotically converges on designing AGI suggests that it is, in fact, "deep tech". And this property has to mesh with other practical constraints on database behavior. Many problems in databases are hard with little forward progress in decades.
It is true that most database research is not deep tech but there is ample room for it to be if one is sufficiently ambitious.
The dude doesn't know anything about databases.
I'm partly kidding; there's plenty of space for innovation in databases. You should sponsor https://sled.rs/
This is essential for a company that wants to stay on the frontier of database tech.
Performance is a feature. "Written in C++" is a strange idea of a feature.