Naively, you might pick a single set of benchmarks to measure all historical models, but this is a mistake.
方法论上最值得借鉴的一句:把 LLM 史切成三个时代、各用当代基准分别打分,避开了饱和基准把差距压平的陷阱。代价是三段分数彼此不可比,所以"追赶期减半"是趋势描述,不是能往外推的定律。
Naively, you might pick a single set of benchmarks to measure all historical models, but this is a mistake.
方法论上最值得借鉴的一句:把 LLM 史切成三个时代、各用当代基准分别打分,避开了饱和基准把差距压平的陷阱。代价是三段分数彼此不可比,所以"追赶期减半"是趋势描述,不是能往外推的定律。
Pulled the trigger today & switched 100% of Lindy traffic to DeepSeek v4, churning from Anthropic models. Saves us millions of $ & we're actually seeing an _increase_ in performance on many core use cases
与行业普遍认为闭源模型性能优于开源模型的认知相反,Lindy的案例显示切换到开源模型不仅节省大量成本,还提高了性能,这一发现挑战了闭源模型优越性的主流观念。
(1:21:20-1:39:40) Chris Aldrich describes his hypothes.is to Zettelkasten workflow. Prevents Collector's Fallacy, still allows to collect a lot. Open Bucket vs. Closed Bucket. Aldrich mentions he uses a common place book using hypothes.is which is where all his interesting highlights and annotations go to, unfiltered, but adequately tagged. This allows him to easily find his material whenever necessary in the future. These are digital. Then the best of the best material that he's interested in and works with (in a project or writing sense?) will go into his Zettelkasten and become fully fledged. This allows to maintain a high gold to mud (signal to noise) ratio for the Zettelkasten. In addition, Aldrich mentions that his ZK is more of his own thoughts and reflections whilst the commonplace book is more of other people's thoughts.