"We aggregate the user's history in ClickHouse and use it as a data store for training and inference. Even when reading 10s of millions of rows, the performance was very nice and not the bottleneck when training new models."
- Use cases
- Machine learning and GenAI
机器学习与生成式 AI
为机器学习工作负载提供动力的终极实时数据库。借助 ClickHouse,在分析数据中释放生成式 AI 能力从未如此简单。
- 免除专用 ML 数据存储,简化您的数据栈
- 使用极速聚合进行数据准备,支撑 PB 级模型训练
- 通过线性与近似检索技术,执行高速且高效的向量检索
- 即插即用地接入任意厂商的预构建模型
- 通过丰富的集成体系,继续使用您熟悉的 ML 工具进行开发
了解为什么众多企业选择 ClickHouse 驱动其 AI 工作负载。
业界领先的摄取速率,可处理持续不断的数据流,让您始终基于最新信息做出准确预测与决策。
大规模下无与伦比的查询性能。毫秒级查询数十亿行数据,缩短迭代周期,最大限度地提升数据效率。
强大的自动扩缩能力,专为应对不可预测的工作负载而生。专注于机器学习,无需操心底层基础设施。
可作为 Python 内嵌式 OLAP SQL 引擎 使用。通过 chDB,在 Python 代码中直接释放 ClickHouse 的全部能力。
赢得众多 大规模 数据开发者的信赖
ClickHouse 助力 ML 与 AI
ClickHouse 为从复杂数据中轻松获取洞察而生,无论数据量多大。无论您是通过聚合提取模型训练与评估所需的关键信息,通过用户自定义函数运行推理,还是执行向量检索,ClickHouse 都能让您最大限度地提升数据效率,为任何应用释放 AI 的力量。
ClickHouse is trusted at scale to ingest and process billions of new events per day from a wide range of sources and formats. For continuous streams of data, ClickPipes seamlessly manages your ingestion pipelines so that you don't have to.
Features like User Defined Functions, described in more depth below, can be used to invoke models at insert time. This gives you the ability to pass incoming data to a model, receive the output, and store these results along with your ingested data. All without having to spin up other processes or jobs.
Native table functions make it easy to query data wherever it lives, whether locally or in object stores such as GCS and S3, or applying transformations via services like HuggingFace.
ClickHouse User Defined Functions give you the flexibility to run Python scripts - or whichever executable language you prefer - directly in ClickHouse. These scripts can be triggered at insert or query time, making it easy to invoke pre-built models from providers like OpenAI and HuggingFace, or your own.
Our extensive suite of statistical and aggregation functions scale seamlessly over petabytes of data, providing powerful model training and evaluation resources. With support for the most granular precision data types and codecs, you don't need to worry about reducing granularity.
With ClickHouse, executing vector searches using linear or approximate techniques is effortless, with out-of-the-box support and blazing speed.
ClickHouse is trusted all over the world to power customer-facing applications, where real-time responsiveness is critical.
With ClickHouse, you have everything you need to enrich your customer experiences through machine learning workloads run on your data, all in one place.
Our vibrant and growing ecosystem of integrations makes it easy to leverage your notebooks, visualization tools, and more, directly with ClickHouse.
打造有价值的体验与洞察
无论您是在构建引人入胜的个性化功能、在产品中融入语义搜索,还是自动从原始内容中生成摘要洞察,ClickHouse 都能为您提供构建数据驱动的 AI 功能所需的能力。
统一您的数据栈
无需为特定机器学习任务(如向量检索)单独部署专用数据存储。借助 ClickHouse,您可以使用统一的数据存储,一站式完成分析支撑、运行机器学习工作负载并完成临时查询。
高效管理数据
ClickHouse 凭借高效的资源管理大幅提升成本效益。列式存储设计带来业界领先的压缩比,显著降低存储压力,并为最繁重的 ML 工作负载保持极致速度。
使用您喜爱的工具
直接将 ClickHouse 与您喜爱的 ML 工具一同使用。我们持续壮大的集成生态涵盖主流机器学习框架、可视化工具、Notebook 等。
Supporting references
如需了解如何在 ML 中入门使用 ClickHouse 的详细指南,请查阅我们的博客:
- Vector Search with ClickHouse - Part 1
- Vector Search with ClickHouse - Part 2
- Video: ClickHouse for AI - Vectors, Embedding, Semantic Search, and more - Alexey Milovidov, ClickHouse
- Video: Vector Search In ClickHouse - Dale McDiarmid
- Using Langchain with ClickHouse
- Using Deepnote with ClickHouse
- Analyzing Hugging Face datasets with ClickHouse
- Using ClickHouse UDFs to integrate with OpenAI models
- Forecasting Using ClickHouse Machine Learning Functions
- Helicone's Migration from Postgres to ClickHouse for Advanced LLM Monitoring
- ClickHouse and the Machine Learning Data Layer
- Powering Feature Stores with ClickHouse
免费开始使用 ClickHouse Cloud
我们将为您提供 30 天试用期及 300 美元额度,助您自如开启探索。