夺攒
河浪:布上的水渍
各楞:固体搅拌
豁楞:液体搅拌
缺德加冒烟
你消停的啊,一点儿老实气都没有
吵吵巴火的
班大班的
夺攒
河浪:布上的水渍
各楞:固体搅拌
豁楞:液体搅拌
缺德加冒烟
你消停的啊,一点儿老实气都没有
吵吵巴火的
班大班的
如何快速查找使用示例
比如某个spark config,可以用spark.sql.shuffle.partitions + github 的方式搜索
https://github.com/imarvinle/awesome-cs-books
| 1 | Safari账号 | 知名技术出版社O’Reilly运营的电子书平台,有超过4万种技术和管理英文电子,还有众多学习视频等 | https://www.safaribooksonline.com/home/ | Safari:账号:yizhang37@acm.orgACM登录信息:账号:library1@meituan.com密码:mit12345 | Safari操作手册 |
| 2 | ACM Library | 学术期刊 | https://dl.acm.org/ | 账号:library1@meituan.com密码:mit123456 | ACM Library使用手册 |
| 3 | IEEE CS Library | 学术期刊 | https://www.computer.org/csdl | 账号:library1@meituan.com密码:mit12345 | |
| 4 | CCF会员 | 中国计算机学会的期刊、视频资源 | http://dl.ccf.org.cn/index.html | 账号:library1@meituan.com密码:mit12345到2018年12月31日 | |
| 5 | 哈佛商业评论网站 | 订阅:http://shop.caijingmobile.com/product/view/id/341网站:http://www.hbrchina.org/ | 账号:library1@meituan.com 密码:mit123452019.06.14 - 2020.06.14 | ||
| 6 | 知网 | 学术期刊 | http://www.cnki.net/ | 账号:library1@meituan.com密码:mitmitmit没有时长和会员限制,充值,0.5元/页 | |
| 7 | 极客时间 | 在线专栏 | https://km.sankuai.com/page/109336610 | 账号1:18210016864密码1:mit12345到2019年08月04日账号2:18612256271密码2:mit12345到2020年02月19日为了避免被踢,可以使用极客时间小程序,亲测不会被踢~操作步骤:微信搜索【极客时间】小程序 - 我 - 登录 | 07 极客时间使用手册 |
| 8 | 华章电子书库 | 电子书及在线课程 | 华章电子书使用手册 |
https://0x0fff.com/spark-architecture/
https://blog.cloudera.com/how-to-tune-your-apache-spark-jobs-part-2/
腾讯视频
streaming流程中 partition or task数量是如何变化的
kafka offset如何查看
kafka earlist 真的是从最早的吗,还是从被消费掉的数据开始
kafka消费组问题
怎么验证eventhub在不同consumer group中是分别获取的数据
eventhub jar包使用报错,虽然打包到application jar里
spark 程序中的class 为什么需要序列化
sparkSession sparkContext sparkContext中不包含sparkSession的config信息
将kafkaProducer broadcast 遇到 can’t serilize lamda 的问题
2021-08-19T15:34:18.109Z 型字符串的理解与解析
spark提交任务需要时间, 间隔短的batch任务就不适合,会受到严重影响
同样的任务,根据参数选择不同的输入路径,将数据传输到kafka,其中a
a的任务,基本可以保证每分钟提交,b、c都要3-4分钟提交一次,a的量比线上要大几倍(但这可能是写入到cosmos数据就多了的原因),可是数据量大,还可以很快的提交
Thread.sleep(sleepMillisecond) 在driver上执行有什么影响
为什么protobuf解析出问题了
为什么读eventhub有receiver问题
任务执行任务过长,导致平均写入的文件少,无法复现throttling
spark streaming是如何确认开始处理一批数据的? 使用线上的数据每批就很少,测试数据每批就很大。当然测试数据没有线上数据输入的平滑
driver的日志为什么只打印出了系统的info,程序里的info没有打印
怎么查看azure的data center
20 structured streaming partition executor task executor-core 关系
https://blog.csdn.net/mzqadl/article/details/104217828
21.spark写文件的request请求
org.apache.kafka.common.errors.TimeoutException: Expiring for has passed since batch creation
spark.locality.wait
repartition() vs coalesce()
https://stackoverflow.com/questions/31610971/spark-repartition-vs-coalesce
How to optimize number of executor instances in spark structured streaming app?
sleep 是否会占用CPU
https://blog.csdn.net/weixin_41960204/article/details/106785986
https://www.cnblogs.com/yu6688/p/14443104.html
structured streaming 与 kafka
https://spark.apache.org/docs/2.3.0/structured-streaming-kafka-integration.html
https://www.cnblogs.com/yyy-blog/p/12753924.html
解决问题方法论
对问题列checklist
https://github.com/Azure/azure-event-hubs-for-kafka/issues/35
Hi Patrick Chen, we have outputted several hours of data to cosmos, which were based on historical data in 2021/10/13. You may use it temporarily. The online data will be ready soon.
[toc]
Virtual Cluster: a management tool
– Allocates resources across groups within OSD
– Cost model captured in a queue of work (with priority) within the VC
One Incident Management System for Microsoft.
类似于share的功能
(Data Warehouse Controller) is an orchestration system built by Ads Data team, aiming at schedule tasks with a dependency system.
S2B streaming to batch
Structured Computation Optimized for Parallel Execution
This module is the second of a three-part training series that provides you with insights into how to recognize security threats, helps you adopt healthy security behaviors, and gives you information on the tools and resources available