[{"content":"This blog was born on 2021-10-16. The original site was at blog-hexo. Since all static files were already compiled, no further modifications were made; the subsequent migration should be completed gradually.\nThe reason is that I messed up (sb operation), and the original npm environment could not be restored. Fortunately, I had previously installed Hugo, so this blog\u0026rsquo;s deployment is now handled by Hugo + GitHub Action.\nApart from this site introduction, articles are categorized into four types based on their primary writing purpose, with specific topics, methods, and tool usage organized via tags.\nEnvironment-related configurations and some basic notes will be placed in my Knowledge Base, while some summary articles that I have organized into systematic frameworks will be posted here. Therefore, the update frequency of this blog will be relatively low. When the knowledge base reaches a certain level and can be archived and organized, those contents will appear in the form of blog posts. Most of the content actually carried by the blog is daily updates \u0026ndash;.\nblog: Machine learning, mathematical analysis, programming practice, project development, and tool usage. essays: Personal opinions, miscellaneous notes on daily life, and reflections on tools, communities, and life. journey: Academic pursuits, exams, competitions, graduation farewells, and annual reviews. note: Book reviews, reading reflections, and thoughts during reading. Then, I continued using MathJax instead of KaTeX for formula rendering. One reason is to facilitate readers in exporting formulas; the other is that I encountered issues with inline matrix rendering when using KaTeX, which I couldn\u0026rsquo;t resolve\u0026hellip; So now it\u0026rsquo;s actually in the form of MathJax + short code.\nAdditionally, after a lot of tinkering, this blog has not only some color optimizations but also added a friends link page, and even supports sass and scss\u0026hellip;\nBelow is a test of hint\u0026rsquo;s short code\ninfo\nThis is a test for hint info note\nThis is a test for hint note important\nThis is a test for hint important warning\nThis is a test for hint warning danger\nThis is a test for hint danger tip\nThis is a test for hint tip example\nThis is a test for hint example ","permalink":"https://blog.bj-yan.top/en/p/readme/","summary":"\u003cp\u003eThis blog was born on 2021-10-16. The original site was at \u003ca href=\"https://blog-hexo.bj-yan.top/\"\u003eblog-hexo\u003c/a\u003e. Since all static files were already compiled, no further modifications were made; \u003cdel\u003ethe subsequent migration should be completed gradually.\u003c/del\u003e\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eThe reason is that I messed up (sb operation), and the original \u003ccode\u003enpm\u003c/code\u003e environment could not be restored. Fortunately, I had previously installed \u003ccode\u003eHugo\u003c/code\u003e, so this blog\u0026rsquo;s deployment is now handled by \u003ccode\u003eHugo + GitHub Action\u003c/code\u003e.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eApart from this site introduction, articles are categorized into four types based on their primary writing purpose, with specific topics, methods, and tool usage organized via tags.\u003c/p\u003e","title":"About This Blog: Categories, Tech Stack, and Writing Plans"},{"content":"Preface Open a craft beer bar\u0026rsquo;s menu, and you\u0026rsquo;ll often see a string of abbreviations: IPA, NEIPA, DIPA, ABV, IBU\u0026hellip; The Chinese names aren\u0026rsquo;t much easier either: Pilsner, Wheat, Saison, Porter, Stout, Imperial, Barrel-aged. Before even taking a sip, it feels like you\u0026rsquo;re about to tackle a reading comprehension test.\nHowever, you don\u0026rsquo;t need to memorize every classification before drinking beer. Knowing these names helps you find flavors you enjoy and makes your next order less of a gamble. This post starts with the most basic concepts to compile an entry-level guide that is both understandable and practical.\nWhat Exactly Is \u0026lsquo;Craft Beer\u0026rsquo;? The Chinese term \u0026lsquo;jingniang\u0026rsquo; corresponds to Craft Beer, but it is not a specific beer style, nor is there a single definition used universally worldwide. The Brewers Association\u0026rsquo;s definition specifically targets American craft breweries, emphasizing small scale and independent ownership; this is an industry statistical metric and cannot be directly applied as a standard for judging all beers in all countries. 1\nIn everyday conversation, \u0026lsquo;craft\u0026rsquo; more often evokes exploration of flavor and recipes: different malts, hops, and yeasts, as well as changes brought by fruits, spices, or barrel aging. But the following judgments are unreliable:\nCraft does not equal IPA. Craft breweries also brew Lagers, Wheat beers, and other traditional styles. Craft does not equal bitter, strong, or high alcohol. Refreshing, low-bitterness beers can be equally interesting. Hazy, unfiltered, or additive-free are not sufficient conditions. Clear beers can be delicious; using adjuncts reasonably does not imply poor quality. Price and packaging cannot substitute for taste. No matter how beautiful the label, what matters is the beer in the glass. Instead of arguing whether a bottle is \u0026lsquo;craft\u0026rsquo; enough, ask: What style is it? Is it fresh? Are the flavors balanced? Do I like it?\nWhere Does a Beer\u0026rsquo;s Flavor Come From? Four Basic Ingredients Beer is typically understood through water, malt, hops, and yeast, though ungerminated grains, sugar, fruits, and spices may also be used.\nIngredient Primary Role Flavor or Mouthfeel to Watch For Water Provides the brewing base; mineral composition affects brewing and flavor expression The same recipe can yield different levels of smoothness and bitterness Malt Provides starch for mashing and influences color and flavor Grain, bread, biscuit, caramel; darker malts may also impart coffee or roasted notes Hops Provides bitterness, as well as various aromas and flavors Floral, herbal, citrus, pine, or tropical fruit notes Yeast Converts fermentable sugars into alcohol and carbon dioxide, and creates fermentation flavors Some yeasts impart fruity or spicy notes, while others produce cleaner profiles Therefore, a fruity aroma doesn\u0026rsquo;t necessarily mean fruit was added; a coffee taste doesn\u0026rsquo;t necessarily mean coffee was added. Only the ingredient list can confirm whether these were truly included. 2\nSimplified Understanding of the Brewing Process After crushing the malt, it is mixed with water and undergoes mashing to convert starches into sugars. The wort is separated, boiled, and hops are added according to the recipe. Then it is cooled, inoculated with yeast, fermented, aged, and packaged before becoming the beer we drink.\nAdding hops at different stages yields different effects. Hops added during the boil help develop bitterness; those added later focus more on aroma. Common dry hopping involves adding hops during or after fermentation, primarily to enhance hop aroma, not literally \u0026lsquo;drying the hops before adding them\u0026rsquo;. 2\nFirst, Distinguish Ale from Lager Ale and Lager are primarily fermentation classifications, not quality grades. For beginners, think of it this way: Ales typically use ale yeast fermented at higher temperatures, while Lagers use lager yeast fermented at lower temperatures and undergo cold conditioning. In actual brewing, exceptions exist, so conclusions shouldn\u0026rsquo;t be drawn solely based on temperature or taste.\nAles include Wheat, IPA, Porter, and Stout; Lagers include Pilsner, Munich Helles, Dark Lager, and Bock. Ales aren\u0026rsquo;t necessarily heavy, and Lagers aren\u0026rsquo;t necessarily light. \u0026lsquo;Top-fermenting\u0026rsquo; and \u0026lsquo;bottom-fermenting\u0026rsquo; are common terms, but don\u0026rsquo;t interpret them as yeast working only in one layer of the liquid. 3\nDecoding the Label: Common Terms How to Read the Numbers? Label Meaning How It Helps with Ordering ABV (Alcohol by Volume) Alcohol by Volume For example, 5% ABV means alcohol makes up about 5% of the total volume. Combined with serving size, it determines total alcohol intake IBU (International Bitterness Units) International Bitterness Units Describes the level of bittering compounds, but does not equal perceived bitterness; sweetness, body, and other factors also influence perception OG / FG (Original / Final Gravity) Original Gravity / Final Gravity Correspond to the specific gravity of wort before fermentation and beer after fermentation, commonly used in brewing logs and alcohol estimation °P (Degrees Plato) Degrees Plato Often used on labels to indicate original extract concentration, reflecting the extract content of wort before fermentation, not alcohol content SRM / EBC Two color scales Higher values generally mean darker color; the two scales cannot be used interchangeably; darker color does not imply higher alcohol content For most consumers, checking the style, ABV, and packaging date is sufficient. IBU can help inform your choice, but there is no need to treat it as a flavor ranking list. 2\nHow to read flavor and process? Body: How light or full the beer feels in the mouth; it is not about the beer\u0026rsquo;s color. Dry: Usually describes lower sweetness and a clean finish; it does not mean bitter, nor does it imply the complete absence of sugar. Esters and Phenols: In flavor descriptions, esters often evoke fruit, while phenols may suggest spices like cloves. Whether they are appropriate depends on the specific style. Carbonation: The prickly sensation caused by carbon dioxide. It is a different dimension from body and smoothness. Finish / Aftertaste: The sensation left after swallowing, such as roasted notes, bitterness, or dryness. Barrel-aged: Aged in wooden barrels, which may introduce wood flavors or flavors from the wine or spirits previously held in the barrel; it is not an independent base style. Bottle-conditioned: Fermentation continues in the bottle after packaging, usually to create carbonation, and may leave yeast sediment. 2 There are also terms often found in beer names: Session usually emphasizes lower alcohol and drinkability; Double / Imperial typically indicates a stronger version; DDH (Double Dry Hopped) refers to double dry hopping, but the specific number of additions, quantities, and naming conventions vary by brewery. These names are not a unified ruler; especially, \u0026ldquo;Double\u0026rdquo; does not mean the alcohol content is exactly doubled. 4, Dry hopping reference 2\nCommon styles: How do they taste different? Below, common styles are introduced by flavor direction. They serve as signposts to help understanding, not templates that every bottle must strictly follow; even within the same style, there can be significant recipe differences.\nLagers and Pilsners: Crisp with layers Lager is a broad category, and Pilsner (Pilsner / Pils) is a specific style within it. Pilsners often have a more distinct hop bitterness and expressions of herbal or floral notes on top of crispness; Munich Helles leans more toward a soft malt character. Dark lagers also exist, so color alone cannot determine if a beer is a lager.\nWhen drinking these, pay attention to whether the malt and bitterness are balanced and whether the finish is clean. \u0026ldquo;Light\u0026rdquo; may simply mean the flavors are not assertive, not that there is no brewing skill involved. 3\nWheat beers: German and Belgian styles differ German Wheat Beer (Weissbier / Hefeweizen) commonly features fermentation aromas reminiscent of banana and cloves, with rich foam and low bitterness. The banana flavor here usually comes from the yeast, not from adding actual bananas. 5\nBelgian Witbier also uses wheat but is often paired with spices like coriander seeds and orange peel, resulting in a generally crisp profile with citrus and spice notes. The \u0026ldquo;white\u0026rdquo; in the name refers to the style, not the actual color of the liquid. 6\nIf you find standard lagers a bit monotonous but do not want to immediately face strong bitterness, both of these directions are worth exploring. When ordering, specifying German wheat beer or Belgian witbier is more accurate than simply saying \u0026ldquo;wheat beer.\u0026rdquo;\nPale Ale: An intermediate stop for understanding hops Pale Ale is not a fixed recipe. Common American Pale Ales (APA) emphasize hop aroma while retaining malt support; compared to American IPAs, they are typically more restrained in strength and hop expression.\nIf you want to understand what \u0026ldquo;hop aroma\u0026rdquo; is but worry that IPA is too intense, you can start with Pale Ale. Here, \u0026ldquo;Pale\u0026rdquo; is relative to traditional dark beers and does not guarantee that every pour will be particularly light in color. 7\nIPA: Not just one kind of \u0026ldquo;bitter beer\u0026rdquo; IPA stands for India Pale Ale. Today\u0026rsquo;s IPA has branched into many directions; just looking at these three letters cannot accurately predict the taste in the glass.\nAmerican IPA / West Coast IPA: Emphasizes hops, with possible expressions of citrus, pine resin, floral notes, etc. The typical West Coast direction is usually drier with more distinct bitterness. If you enjoy grapefruit-like bitterness and a clean finish, look in this direction. 7\nHazy IPA / New England IPA (NEIPA): Emphasizes tropical fruit-like hop aromas and a soft, full mouthfeel; perceived bitterness is usually more restrained than in traditional IPAs. \u0026ldquo;Juiciness\u0026rdquo; does not mean juice was added, nor does it imply the absence of alcohol or bitterness; haze level is not a quality score. 8\nDouble IPA (DIPA) / Imperial IPA: Typically have higher alcohol content and stronger hop expression. Liking a standard IPA does not guarantee you will enjoy the stronger version; alcohol presence, body, and drinking rhythm will all change. 3\nWhen ordering, saying \u0026ldquo;I want a beer with prominent hop aroma, but not too bitter, and not too high in alcohol\u0026rdquo; is often more helpful than saying \u0026ldquo;Give me an IPA.\u0026rdquo;\nPorters and Stouts: Coffee, Chocolate, and Roasted Notes Porters and Stouts are both common dark style families, often featuring roasted, coffee-like, or cocoa-like flavors. However, in modern recipes, they overlap, and you cannot draw an absolute boundary between them based solely on color or a single ingredient. 3\nTake Irish Stout as an example: it can have a distinct roasted character and a relatively dry finish, yet not necessarily have a high alcohol content. Black does not equal strong, and a thick head does not equal sweet. 9\nSome stouts incorporate ingredients like lactose, coffee, cocoa, or vanilla, moving toward a sweeter, dessert-like expression. When you see \u0026ldquo;Milk Stout\u0026rdquo; or \u0026ldquo;Milk Stout,\u0026rdquo; check the ingredients: lactose is not always fully fermented by standard beer yeast, so those with lactose intolerance should pay special attention. 2\nIf you enjoy coffee or dark chocolate, you can start by exploring roasted flavors; but be sure to distinguish whether you prefer a dry roastiness or a rich, dessert-like mouthfeel.\nBelgian Styles: Let the Yeast Take Center Stage Belgian styles are not a single flavor profile. Take Tripel as an example: it is typically pale in color but has a high alcohol content, along with fruity and spicy notes from fermentation and a relatively dry finish. The \u0026ldquo;three\u0026rdquo; in the name does not mean three fermentations, nor does it represent triple the alcohol content. 10\nOther common styles include the darker, malt-forward Dubbel, and the typically dry, yeast-aromatic Saison. These names describe different traditional styles, not different sizes or strength levels of the same beer. 3\nThese beers remind us that beer aroma doesn\u0026rsquo;t come only from hops. When you smell spices or fruit, it\u0026rsquo;s worth thinking about the yeast first.\nSour Beers and Fruit Beers: Distinguish Sour, Sweet, and Fruity Notes Sour beers are a group of beers with diverse flavors and brewing methods, not a fixed recipe. For example, Gose typically has a sour note, restrained saltiness, and coriander character; it is not \u0026ldquo;fruit soda with salt added.\u0026rdquo;11\nFruit beers can be built on various base styles, and they may or may not be sour. \u0026ldquo;Fruit sour\u0026rdquo; and \u0026ldquo;fruit puree / smoothie\u0026rdquo; on a menu cannot be directly equated: the latter often describes products with more fruit puree and a thicker mouthfeel, but naming conventions vary across breweries. 3\nIf you enjoy fruity aromas, it\u0026rsquo;s best to confirm when ordering: How high is the acidity? Is it sweet? Is the mouthfeel crisp or thick? Just naming your favorite fruit may still result in a beer that is completely different from your expectations.\nHow to Order at a Craft Beer Bar for the First Time Start by describing your preferences, then choose a style. Below are suggested directions based on the flavor characteristics discussed earlier; these are not guaranteed recommendations for everyone:\nWhat you want to drink Directions to explore first Add this when ordering Crisp, clean, not too heavy Helles, Pilsner A bit of bitterness is okay, but keep it low? Aromatic, but not too bitter German Wheat Beer, Belgian Witbier Do you prefer banana and clove, or citrus and spice? Want hop aroma APA, Hazy IPA Prefer lower alcohol, and confirm the bitterness level Like dryness and distinct bitterness West Coast IPA How strong a bitterness and alcohol presence can you accept? Like coffee or cocoa Porter, Stout Do you prefer it dry, or sweet and full-bodied? Like sourness or fruit Sour Beer, Fruit Beer Confirm acidity, sweetness, and fruit puree presence If the bar offers small pours or tasting flights, choose a few distinctly different styles to compare. There\u0026rsquo;s no need to try every style on your first visit, and you shouldn\u0026rsquo;t force yourself to accept flavors you dislike just because it\u0026rsquo;s your \u0026ldquo;beginner\u0026rdquo; experience.\nRather than memorizing beer names, it\u0026rsquo;s more practical to note three pieces of information: style, ABV, and the reason you liked or disliked it. For example, \u0026ldquo;I like the citrus aroma, but the finish is too bitter\u0026rdquo; gives you a clear direction for next time.\nHow to Drink to Better Distinguish Flavors You don\u0026rsquo;t need a complex ritual. Use a clean glass free of detergent residue, pour a moderate amount, and observe in the following order:\nLook: Color, clarity, and foam. These help identify the style but cannot determine quality on their own. Smell: First, identify the most obvious impression: bread, citrus, banana, or coffee? No need to write down ten adjectives at once. Taste: Pay attention to sweetness, bitterness, sourness, body, carbonation, and alcohol presence. Wait: Focus on the finish and aftertaste, then decide if you want another sip. Temperature also changes the experience. When too cold, aromas and flavors are often harder to distinguish; crisp styles can start at lower temperatures, while rich styles can be compared as they naturally warm in the glass. Prioritize the brewery\u0026rsquo;s serving recommendations; there\u0026rsquo;s no need to chase a single \u0026ldquo;standard temperature\u0026rdquo; for all beers. 2\nThere is no standard answer for describing taste. Saying \u0026ldquo;I think it tastes like grapefruit peel,\u0026rdquo; \u0026ldquo;It\u0026rsquo;s too bitter after drinking,\u0026rdquo; or \u0026ldquo;This reminds me of toasted bread\u0026rdquo; is more valuable than forcing a string of technical terms.\nStorage and Common Misconceptions After buying, check the storage instructions on the label, especially for products requiring refrigeration. Beers with prominent hop aroma are usually more sensitive to freshness; fruit puree products that require refrigeration must follow the brewery\u0026rsquo;s guidelines. Avoid prolonged exposure to high temperatures or direct sunlight, and don\u0026rsquo;t automatically assume \u0026ldquo;craft beer\u0026rdquo; means it\u0026rsquo;s suitable for aging. 2\nFinally, here are a few commonly confused terms:\n“Original gravity of 12° means an alcohol content of 12%.” These are two different metrics; alcohol content must be checked via ABV. “If IBU doubles, it will definitely taste twice as bitter.” Perceived bitterness is also influenced by sweetness, body, and overall balance. “A banana aroma means bananas were added.” The characteristic banana note in German wheat beers can come from yeast fermentation. “Dark beer always means high alcohol.” Dark ingredients affect color, but they cannot substitute for alcohol content labeling. “The cloudier, the better.” Some styles should be cloudy, others clear; judgment should be based on the specific style. “If it tastes fruity like a soda, I can drink it freely.” Easy drinkability doesn\u0026rsquo;t change the alcohol content; you still need to watch the ABV and actual consumption. This discussion focuses on flavor and drinking experience, not on providing health advice regarding alcohol. Do not drive after drinking; not drinking at all does not prevent you from understanding these concepts.\nConclusion The most interesting thing about craft beer is that while they are all called \u0026ldquo;beer,\u0026rdquo; they can lead to vastly different flavors: from fresh malt notes, to banana and clove, citrus and pine resin, to roasted and sour profiles.\nTerminology and classification help us communicate more easily, but ultimately there is no need to turn drinking into an exam. Being able to articulate what you like and dislike, and choosing your next beer with more intention, is where this introductory knowledge becomes useful.\nFurther Reading Brewers Association: Definition of American Craft Breweries—Understand industry definitions and scope of application.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCraftBeer.com: Beer Glossary—Look up ingredients, processes, and tasting terminology.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBJCP 2021 Beer Style Guidelines—Query aroma, flavor, mouthfeel, and parameter ranges for specific styles; it primarily serves judges and communication, not as a scoring sheet for personal preference.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNaming and strength reference \u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGerman wheat beer reference \u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBelgian witbier reference \u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPale Ale vs. IPA comparison \u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHazy IPA reference \u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIrish Stout reference \u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTripel reference \u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGose reference \u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-craft-beer/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eOpen a craft beer bar\u0026rsquo;s menu, and you\u0026rsquo;ll often see a string of abbreviations: IPA, NEIPA, DIPA, ABV, IBU\u0026hellip; The Chinese names aren\u0026rsquo;t much easier either: Pilsner, Wheat, Saison, Porter, Stout, Imperial, Barrel-aged. Before even taking a sip, it feels like you\u0026rsquo;re about to tackle a reading comprehension test.\u003c/p\u003e\n\u003cp\u003eHowever, you don\u0026rsquo;t need to memorize every classification before drinking beer. Knowing these names helps you find flavors you enjoy and makes your next order less of a gamble. This post starts with the most basic concepts to compile an entry-level guide that is both understandable and practical.\u003c/p\u003e","title":"Craft Beer"},{"content":"Good and evil, beauty and ugliness, the transient and the eternal, emptiness and reality, survival and destruction. It seems as if everything in this story is black and white, with clear boundaries. Mishima\u0026rsquo;s descriptions make these opposing forces even more concrete; this intense contrast makes the beauty of the Golden Pavilion even more absolute and its destruction even more tragic.\n\u0026ldquo;Inferiority and Jealousy\u0026rdquo;: Both Mizoguchi and Kashiwagi initially believed that others ignored their physical disabilities as if ignoring the meaning of their existence. Mizoguchi sought equal dialogue, while Kashiwagi chose to accept those who recognized him, believing that someone loved the \u0026ldquo;meaning of his existence\u0026rdquo; he possessed, yet overlooked the truly beautiful aspects of himself elsewhere.\nThe Golden Pavilion was also Mizoguchi\u0026rsquo;s shackle—not because of his intense desire for it, but because his desire was too small. He believed since childhood that the emptiness of the Golden Pavilion\u0026rsquo;s beauty was the ultimate goal, and his desire was confined to the Pavilion, becoming his shackle. The reason he always considered the Golden Pavilion beautiful was that he never sought something more beautiful; even after seeing the real Golden Pavilion, the emptiness of its beauty remained in Mizoguchi\u0026rsquo;s heart. Naturally, the beauty of the Golden Pavilion eventually became a catalyst for Mizoguchi\u0026rsquo;s sins.\nI still cannot fully understand Mizoguchi\u0026rsquo;s complete motivation: was it because he felt the Golden Pavilion was his shackle, or because he was jealous of its beauty and the recognition he could not attain, or was it the evil \u0026ldquo;intent to kill\u0026rdquo; that arose in moments of calm? He thought about the Golden Pavilion several times without truly seeing it; thus, the so-called destruction became meaningless. The beauty of the Golden Pavilion exists eternally and cannot be erased by burning. Burning the Pavilion only causes its physical disappearance, yet makes its beauty even more eternal, so it appears to be a meaningless action for Mizoguchi.\nI also dislike the ending of The Temple of the Golden Pavilion. Perhaps it is because I want to adhere to the facts; after all, it is based on real events, and Mishima only added simple imagination, which is still acceptable. Ending with selfish destruction brings Mizoguchi back to being an ordinary person and returns the beauty of the Golden Pavilion to emptiness, just as the story began. If it were me, I might have turned into an empty illusion along with the Golden Pavilion.\n","permalink":"https://blog.bj-yan.top/en/p/note-jin-ge-si/","summary":"\u003cp\u003eGood and evil, beauty and ugliness, the transient and the eternal, emptiness and reality, survival and destruction. It seems as if everything in this story is black and white, with clear boundaries. Mishima\u0026rsquo;s descriptions make these opposing forces even more concrete; this intense contrast makes the beauty of the Golden Pavilion even more absolute and its destruction even more tragic.\u003c/p\u003e\n\u003cp\u003e\u0026ldquo;Inferiority and Jealousy\u0026rdquo;: Both Mizoguchi and Kashiwagi initially believed that others ignored their physical disabilities as if ignoring the meaning of their existence. Mizoguchi sought equal dialogue, while Kashiwagi chose to accept those who recognized him, believing that someone loved the \u0026ldquo;meaning of his existence\u0026rdquo; he possessed, yet overlooked the truly beautiful aspects of himself elsewhere.\u003c/p\u003e","title":"Reflections on The Temple of the Golden Pavilion: Good and Evil, Beauty and Ugliness, and Emptiness"},{"content":"0x00 A Few Words Before Could I actually publish my year-end summary on time this year?\nIf last year\u0026rsquo;s keywords were \u0026lsquo;street photography\u0026rsquo; and \u0026lsquo;premiere screenings\u0026rsquo;, then I think this year\u0026rsquo;s keywords are \u0026lsquo;alcohol\u0026rsquo;, \u0026lsquo;freedom\u0026rsquo;, and \u0026lsquo;anxiety\u0026rsquo;.\n0x01 Alcohol This year, maybe around March to May, I started drinking alcohol. I drank a lot, and my spending on alcohol wasn\u0026rsquo;t small either. It started with having a bit of whiskey at my desk, then buying my own tequila, and later visiting cocktail bars and craft beer bars\u0026hellip; In the end, I found I preferred craft beer bars. One reason is that craft beers offer richer and more diverse flavors, while many cocktail bars serve common drinks; their signature cocktails often don\u0026rsquo;t match my taste, and you usually can\u0026rsquo;t sample them beforehand. So I prefer craft beer. Also, the alcohol content in cocktails is hard to judge. A few times, I felt like I was just drinking soft drinks, almost like drinking beer mixed with water\u0026hellip; The amount of base liquor added depends entirely on the bar\u0026rsquo;s integrity. Craft beer, on the other hand, comes directly from the brewery\u0026rsquo;s barrels, so there\u0026rsquo;s no room for faking it. Another main difference I feel is that cocktail bars are mostly for a few people to meet up and chat, often with everyone sitting at one table. Craft beer bars have a slightly more relaxed atmosphere, with seats closer together, lower alcohol content, and it\u0026rsquo;s easier to strike up conversations, which is quite nice.\nI did get drunk a few times this year, but I\u0026rsquo;ve gained a clearer understanding of my own tolerance. Basically, I now know how much I can drink to feel relaxed, how much makes me excited, how much gives me a buzz, how much leads to a headache the next day, how much makes me vomit, and how much might cause blackouts\u0026hellip; However, this tolerance is dynamic and depends on my condition at the time.\nI even bought a set of cocktail-making tools, but I haven\u0026rsquo;t used them much. Most of the time, I don\u0026rsquo;t drink just to drink, and I rarely drink alone. I mostly find friends to hang out with, wanting to chat and absorb some of their \u0026rsquo;energy\u0026rsquo;.\n0x02 Friends This brings us to the topic: the friends I\u0026rsquo;ve made this year whom I\u0026rsquo;m very close with were all met at bars. The \u0026lsquo;universe center\u0026rsquo; can indeed introduce you to many high-quality friends! I met everyone in the AI4Life group here, and I really like you all!\n0x03 Anxiety Speaking of this issue, I\u0026rsquo;m getting a bit anxious again. The main reason is related to my application process. May was truly the peak of my anxiety, and it was also one of the reasons I went to bars frequently. If I hadn\u0026rsquo;t taken some measures, I might have been waiting for the sun to rise just to fall asleep. Later, why did I stop feeling so anxious? Maybe I just started giving up a bit\u0026hellip;\n0x04 Movies \u0026amp; Premieres This year, I caught up on many classic films. I didn\u0026rsquo;t go to the cinema much, mainly because I didn\u0026rsquo;t attend many premiere screenings. The only one I went to was \u0026lsquo;Sheep Without a Shepherd 3\u0026rsquo;, which happened to be the last episode of the \u0026lsquo;No Spoiler Screening Group\u0026rsquo; in 2024, also marking its 500th episode. Unfortunately, no cast or crew members were present.\nOh, and there\u0026rsquo;s movie recaps. For a while, I was crazy about watching movie recaps because I felt one wasn\u0026rsquo;t enough, so I\u0026rsquo;d watch another, and then another\u0026hellip; It never ended. A single movie recap isn\u0026rsquo;t short either, probably around 20 minutes, so time just flew by. Later, when I started watching the actual movies again, I found recaps less interesting. Mostly, I\u0026rsquo;d watch a recap to see what movie it was about, then go watch the movie myself, or check the recap again if there were parts I didn\u0026rsquo;t understand.\n0x05 Books This year might have been the year I read the most books in recent times. A big credit goes to @Yannan. Yannan recommended many very interesting books to me. I love the unique scent of paper books, the sound of turning pages, and the feeling of reading with headphones on the subway\u0026hellip;\nI also enjoy thinking about the viewpoints in books, critiquing the author\u0026rsquo;s subjective opinions, and pointing out their biases. Some of the books I read this year were quite opinionated.\n0x06 Music Like last year, I don\u0026rsquo;t have a membership on any domestic music platforms, so I can\u0026rsquo;t see any year-end listening summaries. My go-to music apps have always been Bilibili and YouTube. After switching to a Mac, I started using Apple Music, but I still prefer finding playlists on Bilibili and YouTube to listen to.\n0x07 Games Looking at Steam\u0026rsquo;s 2024 year-end review, I spent 75% of my gaming time using a controller hhh. I\u0026rsquo;ve basically become a \u0026lsquo;controller dog\u0026rsquo;. For the first five months, my Apex Legends skills were quite high. Then, I played a bit more in September and October, but by the end, I even quit gaming entirely\u0026hellip; Maybe playing a few games a month now.\n0x08 Stargazing Thinking about it, since moving out, because my location is relatively remote, light pollution is less. One late night, I suddenly noticed how many stars were above my head. After that, I brought over the tripod and camera lens I used in my dorm and started astrophotography. Although APS-C sensors are slightly inferior, they\u0026rsquo;re more than enough for fun. I\u0026rsquo;ve captured Venus approaching the Moon, Orion, the Big Dipper, and the Geminid meteor shower (I even captured a full 3 meteors!!). Later, I also did some stacking using Siril, but I can only say I still have a lot of room for improvement in post-processing hhh.\nGradually, I\u0026rsquo;ve come to love looking up at the night sky. Orion in the winter sky is always the easiest to spot. You can see Mars twinkling with a reddish glow with the naked eye. I love stargazing because the stars are just there. I love photographing stars to see some stars, nebulae, and their unique light and colors that I can\u0026rsquo;t see with my naked eye.\n0x09 My Monthly/Annual Subscription Projects Actually, I\u0026rsquo;ve wanted to write about this for a long time, but I kept putting it off. Taking this opportunity to write my year-end summary, I\u0026rsquo;ll simply take stock of my monthly or annual subscription projects.\nFirst up is the Aliyun Drive membership (168/year). Although Aliyun Drive can be quite frustrating—for instance, it doesn\u0026rsquo;t support sharing ZIP files, and its backup feature often creates redundant copies, wasting a lot of space—I\u0026rsquo;m satisfied with it as a backup cloud drive, especially regarding storage space, file synchronization, and speed. I also like the online playback feature and its potential for extensibility. I\u0026rsquo;ve found files on Aliyun Drive multiple times that I couldn\u0026rsquo;t even locate on my own computer, so I really have to thank its backup function.\nNext is Apple Music: 3 months free, then 11/month. I think it\u0026rsquo;s pretty good; almost all music is searchable, and there aren\u0026rsquo;t many tracks missing due to disputes. Of course, niche music is still somewhat limited.\nThen there\u0026rsquo;s the Fengchao membership. After moving out, I used Fengchao frequently, but the free storage is only 18 hours—not even 24. So if a package arrives right after I leave in the morning, and I go out again in the evening or forget to pick it up, I\u0026rsquo;ll likely exceed the time limit and have to pay. Once, I had quite a few packages and exceeded the limit for three of them at once. At the time, I thought it was no big deal and just paid. Later, I exceeded the limit again, and when I calculated the fees, they were enough to buy a month\u0026rsquo;s membership. So I decided to just get the membership; no more worrying about forgetting to pick up packages.\nShi Pindao\u0026rsquo;s charging support has been intermittent, mainly because their update frequency can\u0026rsquo;t keep up. If I were to donate 10 yuan per month just to watch one charging video, it would be better to save up for three months and watch three videos instead.\nMobile phone SIM card: yes, I got one and switched it to a 9-yuan number-retention plan. As for its purpose, it\u0026rsquo;s mainly to allow me to open a few more accounts.\nOverall, my monthly fixed expenses now add up to 168/12 + 11 + 5 + 10 + 9 = 49. At first, when I hadn\u0026rsquo;t calculated it, I thought it was quite a lot, but now it seems okay?\n0x0A Freedom Doesn\u0026rsquo;t \u0026lsquo;freedom\u0026rsquo; and \u0026lsquo;anxiety\u0026rsquo; seem a bit contradictory? Let\u0026rsquo;s talk about freedom. I included this keyword in my list for this year but never mentioned it until now. There\u0026rsquo;s no such thing as true freedom; only the freedom of one\u0026rsquo;s own heart and soul. I feel free right now. I do what I want to do. I\u0026rsquo;m still young, and I\u0026rsquo;m willing to try anything I want to do or anything I enjoy. If I had to name one thing I do in my life, I wouldn\u0026rsquo;t say \u0026rsquo;to live,\u0026rsquo; but rather \u0026rsquo;to live my own life.\u0026rsquo;\nThere was a time when I wondered about the meaning of life. Later, I realized that the meaning of life is simple: it\u0026rsquo;s \u0026rsquo;to live.\u0026rsquo; As a living being, just \u0026rsquo;living\u0026rsquo; is already great—that is your meaning in this life. So, what I\u0026rsquo;ve actually been searching for isn\u0026rsquo;t the \u0026lsquo;meaning of life,\u0026rsquo; but rather the \u0026lsquo;meaning of my life\u0026rsquo; or the \u0026lsquo;meaning of living.\u0026rsquo; Living authentically is the most important thing. This might be the biggest realization I\u0026rsquo;ve had this year :)\n0x0B 2025 I don\u0026rsquo;t have any special New Year\u0026rsquo;s wishes. I really dislike the word \u0026lsquo;wish,\u0026rsquo; as if they can only be \u0026rsquo;longings\u0026rsquo; or \u0026lsquo;fantasies.\u0026rsquo; So let me just hope for good health and smooth sailing in 2025.\nWishing you a Happy New Year!\n","permalink":"https://blog.bj-yan.top/en/p/journey-annual-summary-2024/","summary":"\u003ch2 id=\"0x00-a-few-words-before\"\u003e0x00 A Few Words Before\u003c/h2\u003e\n\u003cp\u003eCould I actually publish my year-end summary on time this year?\u003c/p\u003e\n\u003cp\u003eIf last year\u0026rsquo;s keywords were \u0026lsquo;street photography\u0026rsquo; and \u0026lsquo;premiere screenings\u0026rsquo;, then I think this year\u0026rsquo;s keywords are \u0026lsquo;alcohol\u0026rsquo;, \u0026lsquo;freedom\u0026rsquo;, and \u0026lsquo;anxiety\u0026rsquo;.\u003c/p\u003e\n\u003ch2 id=\"0x01-alcohol\"\u003e0x01 Alcohol\u003c/h2\u003e\n\u003cp\u003eThis year, maybe around March to May, I started drinking alcohol. I drank a lot, and my spending on alcohol wasn\u0026rsquo;t small either. It started with having a bit of whiskey at my desk, then buying my own tequila, and later visiting cocktail bars and craft beer bars\u0026hellip; In the end, I found I preferred craft beer bars. One reason is that craft beers offer richer and more diverse flavors, while many cocktail bars serve common drinks; their signature cocktails often don\u0026rsquo;t match my taste, and you usually can\u0026rsquo;t sample them beforehand. So I prefer craft beer. Also, the alcohol content in cocktails is hard to judge. A few times, I felt like I was just drinking soft drinks, almost like drinking beer mixed with water\u0026hellip; The amount of base liquor added depends entirely on the bar\u0026rsquo;s integrity. Craft beer, on the other hand, comes directly from the brewery\u0026rsquo;s barrels, so there\u0026rsquo;s no room for faking it. Another main difference I feel is that cocktail bars are mostly for a few people to meet up and chat, often with everyone sitting at one table. Craft beer bars have a slightly more relaxed atmosphere, with seats closer together, lower alcohol content, and it\u0026rsquo;s easier to strike up conversations, which is quite nice.\u003c/p\u003e","title":"Perhaps a 2024 Year-End Summary"},{"content":"I posted this book review on Douban, but I\u0026rsquo;ll post it here again as well; it\u0026rsquo;s been a while since I last updated my blog.\nSince the entire book consists of interview content, the viewpoints are highly subjective, which isn\u0026rsquo;t necessarily a bad thing. The content is somewhat broad, but I really like it. It\u0026rsquo;s like an anthology series: each section sets a theme, and the rest is free rein, with no constraints on form or direction. Of course, being an interview transcript, there is still a certain degree of guidance involved.\nI think the earlier parts regarding the theory of the body and intimate relationships were well-explained, serving as a sort of mini-review. Many of the viewpoints were new to me, and I really enjoyed the first three chapters.\nBut then again, this book is like a clickbait title; actually, only one section discusses that the core of intimate relationships is friendship, or rather, only Part 1 is relevant. However, I fully agree with this content. Any relationship can be viewed as a form of friendship; friendship is all-encompassing. People simply maintain and play different roles within this friendship, allowing it to connect and sustain itself in various forms.\nThe later section on AI was too superficial. Current ChatGPT will not be the final form of AGI. Some of the content mentioned doesn\u0026rsquo;t even require a \u0026lsquo;philosophical worker like him\u0026rsquo; to explore and summarize. During the research process, there have long been studies that describe these problems more concretely and attempt to solve them. For instance, what he calls \u0026lsquo;copying\u0026rsquo; is actually a manifestation of \u0026rsquo;the harmfulness of synthetic data.\u0026rsquo; Describing it in another language just makes it look more like \u0026lsquo;copying.\u0026rsquo; Could this chapter have been written by GPT? Let me add something else: why do I believe current ChatGPT won\u0026rsquo;t be the final form of AGI? The main reason is that regression based on probability is actually unreasonable. RAG makes content generation more reliable, and CoT allows LLMs to extract more information, but neither represents a \u0026lsquo;great truth in simplicity.\u0026rsquo; Therefore, I believe there will be another massive technological evolution to achieve AGI. Until then, let\u0026rsquo;s wait and see; this AGI might not be made by OpenAI.\nLet me add two more points. First, all the talk about ChatGPT ruling the world comes from mouths other than those of AI researchers. Those who truly work in this field do not believe current AI can rule the Earth, because everyone knows that current AI models lack the ability to change the world. Their fear stems from the speed of technological innovation and the unknown. Currently, we have not endowed it with such capabilities. If one day we directly give it the ability to influence the physical world, that would be truly dangerous (perhaps embodied AI?).\nSecondly, we indeed need to be wary of AI usage in creative works. There\u0026rsquo;s a conspiracy theory that if AGI is ever realized, AI will inevitably gradually gain the ability to manipulate the physical world, slowly control some humans to serve it, and then gradually launch large-scale rule-taking measures. If such creative works are widely used and read, they will inevitably affect human cognition and thought. From a certain conspiracy theory perspective, perhaps this is already happening?\nThe part about France was too boring; I didn\u0026rsquo;t read it carefully.\nRegarding the point in the film documentary section that is incorrect: \u0026lsquo;Returning to the starting point of film, it is the most direct, simple, and plain record.\u0026rsquo; It\u0026rsquo;s better to say this is the starting point of documentaries. I believe that for the viewer, film is about tasting a story and experiencing a life. For the director, it is a way to express their own viewpoint. Cinematic language and aesthetic techniques are forms of expression; editing is the same. All of this is done to better express the director\u0026rsquo;s viewpoint and describe the story. Documentaries or pure dialogues certainly don\u0026rsquo;t need too much artistic form; guidance is the most important thing. One cannot just talk aimlessly.\nYou can check out Bing Shu\u0026rsquo;s interviews; he guides the interviewee to think, then lets them speak from the heart, which might also be what Bing Shu wants to say to the audience.\n\u0026lsquo;Essay films\u0026rsquo; are quite interesting. Filming like an essay is another form of presenting an essay.\nRegarding collectibles, I suddenly had an idea: I want to find a set of lifelong collectibles during my life as a form of expression and what I pursue. The meaning of life? Isn\u0026rsquo;t living itself the meaning of life? The meaning of life does not lie within yourself; what we seek is not the meaning of life, but the meaning and value of our own lives.\nRegarding art, it is said to be a way of consuming money. However, those things that are merely whining without cause should not be called \u0026lsquo;art.\u0026rsquo; Art should have expression, whether in thought, form, or reflecting an era (of course, this is discovered by later generations). This is somewhat contradictory.\nI really like Nietzsche\u0026rsquo;s views on fire and wine. Life unfolds around fire and wine: in the fire, vitality, strength, vigor, frenzy, and joy jump together; with the aid of wine, a chaotic, autonomous, and pleasure-filled body dances. \u0026hellip; Once fire and wine meet, they mutually generate and stimulate each other; their respective energies become fuller, drunkenness nourishes drunkenness, and light illuminates light. \u0026hellip; But when wine and fire are driven out, reason, logic, and knowledge dilute the wine of revelry, while temperance, asceticism, and self-denial extinguish the unrestrained fire.\nAlthough I don\u0026rsquo;t know why, I suddenly remembered the alchemists from \u0026lsquo;Battle Through the Heavens.\u0026rsquo; The condition for becoming an alchemist is the Fire attribute plus a trace of the Wood attribute. Because of the Wood attribute, one can detect or sense things. Thus, that trace of Wood attribute, which seems to be eroded and suppressed by fire yet exists within the fire, is the soul, and it is the reason why alchemists are one in ten thousand. It is the same here: this Wood attribute is like a trace of reason within fire and wine. Even when drunk or unrestrained, there must be a basic bottom line. This is rare and precious, and it is the key factor that allows you to enjoy the fire and wine.\nLet\u0026rsquo;s talk about something else. After finishing this book, I noticed the dark green priority seats on the subway. If someone really wants to offer their seat, they will do so wherever they are sitting. Could the design of these seats discourage some people who otherwise would have offered theirs? I see priority seats as a reminder, but perhaps they lead to fewer people offering seats rather than more. What if the people sitting there are already exhausted and just need somewhere to lean back and rest? They too may need someone to give them a seat, whether or not they belong to the elderly, frail, disabled, or young groups. Don\u0026rsquo;t we all need a \u0026lsquo;seat\u0026rsquo;? The book mentions how capitalism makes people reluctant to say \u0026lsquo;I can\u0026rsquo;t\u0026rsquo; or \u0026lsquo;I\u0026rsquo;m not able to.\u0026rsquo; This situation feels similar.\n","permalink":"https://blog.bj-yan.top/en/p/note-qin-mi-guan-xi-de-he-xin-shi-you-yi/","summary":"\u003cp\u003eI posted this book review on Douban, but I\u0026rsquo;ll post it here again as well; it\u0026rsquo;s been a while since I last updated my blog.\u003c/p\u003e\n\u003cp\u003eSince the entire book consists of interview content, the viewpoints are highly subjective, which isn\u0026rsquo;t necessarily a bad thing. The content is somewhat broad, but I really like it. It\u0026rsquo;s like an anthology series: each section sets a theme, and the rest is free rein, with no constraints on form or direction. Of course, being an interview transcript, there is still a certain degree of guidance involved.\u003c/p\u003e","title":"The Core of Intimate Relationships is Friendship"},{"content":"0. I didn\u0026rsquo;t expect to have such a happy day today!\nLet me jot this down briefly; I still want to get some work done later, since I\u0026rsquo;m not even the slightest bit sleepy yet.\n1. This morning, I woke up to see haha asking Chinese Americans to vote for her. Well, haha it is. She realized this a bit too late; no one believes her anymore, haha.\nThen, in the \u0026lsquo;Get Rich\u0026rsquo; group chat, we briefly discussed PhD planning, unlocking more interesting directions. I feel like I could try researching some of these later. I need to push forward with the applications next week.\n2. Later, Shubai and I went to the National Centre for the Performing Arts to watch a play, \u0026lsquo;The Suspect Sherlock.\u0026rsquo; This was probably my first time ever seeing a stage play. Watson\u0026rsquo;s acting was truly top-notch; switching between these roles felt incredibly difficult. Sherlock, on the other hand, felt just average (though he is the director). Overall, it was fantastic. The play ran for a full 135 minutes, with a portion in the latter half involving audience interaction. The only regret I have is that the beginning felt too dense (I\u0026rsquo;m not sure if this was a recreation of the original script or perhaps a tribute), as many character portrayals could have been simplified. Also, I felt Sherlock\u0026rsquo;s dialogue in the second half was excessive, making him seem quite different from the character established at the start. Regarding the plot, anyone familiar with many mysteries could guess that Sherlock would rise again after falling down; it was basically a test for Watson. Combined with the typical Chinese \u0026rsquo;ending,\u0026rsquo; it\u0026rsquo;s definitely a happy one; a villainous Sherlock was never an option. The suspense was just so-so. Finally, what was special was that we attended the very last performance, which included a large group photo at the end. That was so much fun! Also, aside from the inspector, everyone else had microphones attached to their foreheads, which felt quite magical and even a bit distracting. I kept staring at Watson\u0026rsquo;s microphone the whole time.\n3. After the play, returning to the dorm, I was hit by a stifling, humid air. I couldn\u0026rsquo;t stand the lack of ventilation; it was too damp, almost stuffy. It felt like the perfect environment for growing microbes. So, unable to tolerate it, I decided to come back and sleep. On the way back, I chatted with the taxi driver the entire time, and we kept talking for nearly half an hour after I arrived. The driver is from the 80s generation. He said some of my thoughts on life reminded him of his own at the time, and that there were certain stages where he felt he had lost a lot. Although his definition of \u0026lsquo;winning and losing\u0026rsquo; wasn\u0026rsquo;t entirely reasonable, I could basically get his point. Some of what he said actually made sense\u0026hellip; Indeed, different people grasp different ideas, so more communication is needed. However, one thing he said was quite good: \u0026lsquo;Even if you lose on the path to winning, it\u0026rsquo;s better than losing all the time.\u0026rsquo; He also called me \u0026lsquo;drifting aimlessly day and day,\u0026rsquo; emmm, it actually felt like he hit the nail on the head\u0026hellip; There was something else too: life\u0026rsquo;s path doesn\u0026rsquo;t follow the crowd\u0026rsquo;s detours; once you veer off, it\u0026rsquo;s hard to turn back. Yeah, I roughly understood what he meant. Let me just record this briefly, haha.\n","permalink":"https://blog.bj-yan.top/en/p/misc-20241104/","summary":"\u003ch2 id=\"0\"\u003e0.\u003c/h2\u003e\n\u003cp\u003eI didn\u0026rsquo;t expect to have such a happy day today!\u003c/p\u003e\n\u003cp\u003eLet me jot this down briefly; I still want to get some work done later, since I\u0026rsquo;m not even the slightest bit sleepy yet.\u003c/p\u003e\n\u003ch2 id=\"1\"\u003e1.\u003c/h2\u003e\n\u003cp\u003eThis morning, I woke up to see haha asking Chinese Americans to vote for her. Well, haha it is. She realized this a bit too late; no one believes her anymore, haha.\u003c/p\u003e\n\u003cp\u003eThen, in the \u0026lsquo;Get Rich\u0026rsquo; group chat, we briefly discussed PhD planning, unlocking more interesting directions. I feel like I could try researching some of these later. I need to push forward with the applications next week.\u003c/p\u003e","title":"Miscellaneous Notes: 20241104"},{"content":"1. Finally upgraded Hugo to the latest version. I used to think it was a hassle, but now it doesn\u0026rsquo;t seem to be; just tweaking a few variables does the trick.\nI\u0026rsquo;ve also removed the Douban and recent blog activity links from my GitHub homepage. Lately, I might just vent here on the blog. I don\u0026rsquo;t like domestic platforms—Douban, Weibo, Xiaohongshu, Zhihu—they all have their own distinct crowds. Some international social apps make it hard to find friends, so maybe it\u0026rsquo;s better to manage my own data. Perhaps the blog is still my preferred way for now.\nAlthough I\u0026rsquo;ve had thoughts about running a self-media account, I feel this isn\u0026rsquo;t the right moment. My time, energy, experiences, and topics aren\u0026rsquo;t enough to sustain a long-term self-media presence. I hope to find my own style.\n2. I finally feel a bit more relaxed. First, I resolved the thesis proposal and some confusion about applications, then dealt with finding a place to live. Although I have a To-Do List with many items, at least I can tackle them one by one now.\n3. A couple of days ago, I happily got a private room. Yes, I\u0026rsquo;ve moved out. I didn\u0026rsquo;t want to be too constrained, and I won\u0026rsquo;t look for other excuses anymore. Living on my own allows me to be alone and enjoy the process of organizing my things. More importantly, I feel like I can fully control my time, which feels pretty good.\nI hope to find my own rhythm.\n4. A while back, I met a few new friends. Some coincidences and mutual understanding brought us together.\nI\u0026rsquo;ve never experienced the feeling of having friends with whom I can talk about absolutely everything. Later, I realized more and more that there\u0026rsquo;s a reason we came together—it\u0026rsquo;s not just about personality, attitude, values, and so on.\nThis relaxed and joyful atmosphere while being together made me feel like I\u0026rsquo;ve lived again.\n5. I\u0026rsquo;ve never been a decisive person, but I hope to make some decisions quickly based on my own understanding, even if I might regret them later. But I hope it\u0026rsquo;s not blind or passive.\n6. About Beijing Beijing is getting colder, and autumn has arrived. Seeing my junior\u0026rsquo;s WeChat moments reminded me of myself a year ago. At that time, I had my own rhythm: going to the institute on weekdays and exploring the scenery or taking walks on weekends. It was truly the rhythm I liked. After moving to the Environmental Protection Park area, everything got chaotic. I used to go to the institute to work, and when tired, I\u0026rsquo;d grab my camera to take photos. Now, I haven\u0026rsquo;t picked up my camera in a long time, and even if I did, I wouldn\u0026rsquo;t have the mindset to enjoy that pleasure.\nBeijing is full of smog lately, but I checked my WeChat moments and realized I complained about it around the same time last year, just a bit earlier this year. It seems everything is proceeding in an orderly fashion.\nBeijing is just like that; it keeps running on its own, like a colossal entity, while I feel like I can\u0026rsquo;t find a place to settle.\n7. Just venting a bit; life goes on. Come on, do your best.\nRest when you\u0026rsquo;re tired, but don\u0026rsquo;t rest too long~\n","permalink":"https://blog.bj-yan.top/en/p/misc-20241102/","summary":"\u003ch2 id=\"1\"\u003e1.\u003c/h2\u003e\n\u003cp\u003eFinally upgraded Hugo to the latest version. I used to think it was a hassle, but now it doesn\u0026rsquo;t seem to be; just tweaking a few variables does the trick.\u003c/p\u003e\n\u003cp\u003eI\u0026rsquo;ve also removed the Douban and recent blog activity links from my GitHub homepage. Lately, I might just vent here on the blog. I don\u0026rsquo;t like domestic platforms—Douban, Weibo, Xiaohongshu, Zhihu—they all have their own distinct crowds. Some international social apps make it hard to find friends, so maybe it\u0026rsquo;s better to manage my own data. Perhaps the blog is still my preferred way for now.\u003c/p\u003e","title":"Miscellaneous Notes: 20241102"},{"content":"Preface As everyone knows, TikTok is region-locked and cannot be used normally within China. I\u0026rsquo;ve wanted to try out TikTok for a long time but never succeeded until now.\nToday, while browsing GitHub, I came across this method, but the project\u0026rsquo;s README was incomplete, so I wrote this article to document the process.\nPreparation First, you need a US Apple ID and a scientific internet software such as Shadowrocket, Quantumult X, or Surge.\nAdditionally, the following software needs to be installed:\niTunes v12.6.5.3 (requires a lower version, 64-bit | 32-bit) i4Tools (requires connecting your phone via USB cable) iOS Any Version APP Downloader v6.0: 52pojie Thread Configuring the Proxy You can basically follow this GitHub repository README1, but their documentation is a bit messy. I\u0026rsquo;ve organized it slightly here since I only have Shadowrocket, so I\u0026rsquo;ll use it as the example.\ninfo\nNote: If you have previously installed TikTok, please uninstall it before configuring the proxy!!! Open Shadowrocket\nClick the 配置 at the bottom, then click the i behind your active configuration -\u0026gt; HTTPS 解密 -\u0026gt; 开启 HTTPS 解密 -\u0026gt; 生成新的CA证书 -\u0026gt; 安装证书\nOpen your phone\u0026rsquo;s Settings. The first few lines will show 已下载描述文件. Click 安装, enter the password, click 安装, then click 安装, and finally click 完成\nOpen phone Settings -\u0026gt; 通用 -\u0026gt; 关于本机 -\u0026gt; 证书信任设置 -\u0026gt; Locate Shadowrocket and enable Trust\nOpen Shadowrocket, click the 配置 at the bottom -\u0026gt; 模块 -\u0026gt; Add and unlock the regions you want to view Japan\n1 https://raw.githubusercontent.com/Semporia/TikTok-Unlock/master/Shadowrocket/TiKTok-JP.conf Korea\n1 https://raw.githubusercontent.com/Semporia/TikTok-Unlock/master/Shadowrocket/TiKTok-KR.conf United States\n1 https://raw.githubusercontent.com/Semporia/TikTok-Unlock/master/Shadowrocket/TiKTok-US.conf Taiwan\n1 https://raw.githubusercontent.com/Semporia/TikTok-Unlock/master/Shadowrocket/TiKTok-TW.conf Click the 配置 at the bottom, click the i behind your active configuration -\u0026gt; 规则 -\u0026gt; Top-right + icon -\u0026gt; 类型 -\u0026gt; Select RULE-SET -\u0026gt; 策略 -\u0026gt; Select PROXY or other policy you wish to use (usually the proxy server node corresponding to the region) -\u0026gt; Fill in the Rule Set URL text box\n1 https://raw.githubusercontent.com/Semporia/TikTok-Unlock/master/Shadowrocket/TikTok.list This completes the proxy configuration.\nDownloading Older Versions of TikTok First, download and install iTunes, then log in with your US account.\nOpen iOS Any Version APP Downloader v6.0, select 美国, search for TikTok, and right-click 查看历史版本.\nSearch for Historical TikTok Versions\nScroll down to find the 21.1.0 version and right-click to download.\nLocate TikTok Version 21.1.0\nYou will be redirected to an interception page. At this point, search for TikTok in iTunes, find the corresponding app, and click Download.\nDownload TikTok\nThe download might fail at this stage, but the software has successfully intercepted it. You can pause the download, stop the interception, and then restart the download.\nThe default download path for iTunes is C:\\Users\\用户名\\Music\\iTunes\\iTunes Media\\Mobile Applications. Locate the downloaded TikTok 21.1.0.ipa file.\nInstalling TikTok Open 爱思助手, connect your phone to the computer, and click 信任.\nIf you haven\u0026rsquo;t used it before, 爱思助手极速版 will be installed on your phone. You need to open it once; otherwise, installation might fail due to the service not starting.\nClick 应用游戏, import from local files to install, and locate the TikTok 21.1.0.ipa file you just downloaded to complete the installation.\nIf it asks you to select topics of interest upon opening, the installation was successful! Congratulations, you can now use TikTok normally!\nConclusion Launch!\nTikTok! Launch!\nREADME\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-tiktok-unlock/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eAs everyone knows, TikTok is region-locked and cannot be used normally within China. I\u0026rsquo;ve wanted to try out TikTok for a long time but never succeeded until now.\u003c/p\u003e\n\u003cp\u003eToday, while browsing GitHub, I came across this method, but the project\u0026rsquo;s README was incomplete, so I wrote this article to document the process.\u003c/p\u003e\n\u003ch2 id=\"preparation\"\u003ePreparation\u003c/h2\u003e\n\u003cp\u003eFirst, you need a US Apple ID and a scientific internet software such as Shadowrocket, Quantumult X, or Surge.\u003c/p\u003e","title":"Unlocking TikTok Region Switching Without SIM Card Removal"},{"content":"Hi, long time no see. Today is the third day of the Lunar New Year, Happy New Year!\nI\u0026rsquo;ve always wanted to document the fresh start of 2023. I thought about it a lot, but felt there might not be much worth recording. Still, since it\u0026rsquo;s been so long since I last wrote a blog post, I\u0026rsquo;ll just write something casually.\nFirst, let\u0026rsquo;s casually talk about the year-end summaries from various apps. Bilibili is naturally my perfect-attendance champion! My watch time is almost 2000 hours. The year-end summaries from music apps are no longer very useful, since my music player has completely become YouTube and Bilibili. I can find whatever I want to listen to, there are no copyright issues, and I don\u0026rsquo;t pursue top-tier audio quality, so it\u0026rsquo;s just fine (maybe I should create a browser plugin to automatically track playback time?).\nDouban still has some reference value. The happiest thing this year was probably attending premieres at various cinemas in Chaoyang District. I attended about 6 premieres in 2023: \u0026ldquo;All About Nothing\u0026rdquo; (my first premiere, a nice art film), \u0026ldquo;Operation Moscow\u0026rdquo; (the movie was average, but I got to see Andy Lau), \u0026ldquo;Beast\u0026rdquo; (a terrible movie, but I got to see WaWa), \u0026ldquo;The Hunt\u0026rdquo; (probably the best movie I\u0026rsquo;ve seen at a premiere), \u0026ldquo;Silent Record\u0026rdquo; (an art film dressed as a suspense thriller), and \u0026ldquo;Black Earth Without Words\u0026rdquo; (a TV series premiere, I only watched two episodes, the plot hadn\u0026rsquo;t even unfolded yet, and it ended with a terrible cliffhanger - -|||).\nAlthough I didn\u0026rsquo;t see many good movies at premieres, and their Douban ratings were mostly high, it felt quite nice to meet the creators and interact with them in person hhh. My favorite movies this year were probably \u0026ldquo;Oppenheimer\u0026rdquo; and \u0026ldquo;Johnny Keep Walking!\u0026rdquo; (Annual Meeting Can\u0026rsquo;t Stop), but I had to buy tickets for both myself; I couldn\u0026rsquo;t get into the premieres =A=. Also, \u0026ldquo;Crossing the Line\u0026rdquo; was decent, but I guess it compromised too much to pass censorship. Many movies and TV series suffer from this common issue: they insist on a Happy Ending just to pass review. The theme could have been explored more deeply. Is it that the censors don\u0026rsquo;t understand movies??? Or are they just afraid of getting involved???\nBack in 2022, I was exceptionally active on GitHub, but in 2023 I seemed to have sunk to the bottom of the ocean. Almost all my contributions went to private repositories. The rapid development of LLMs also left me a bit confused. The constant \u0026ldquo;emergence\u0026rdquo; of new tasks in my new job gave me a sense of endless learning hhh. However, most of my work in 2023 had little to do with LLMs. I hope I can work on MM or LMM projects in 2024.\nAt the beginning of 2023, right after the pandemic restrictions were lifted, Beijing was hit hard by the outbreak, so I slipped away early and stayed in Hainan for a month to finish my finals. When I returned, I thought I\u0026rsquo;d be able to go out more now that restrictions were lifted, but I couldn\u0026rsquo;t even get out of Huairou. As expected, Hebei is Hebei; don\u0026rsquo;t get too close to Beijing (x). At the end of the spring semester, I signed up for summer courses early, finished them, and then ran off to Hainan again. Upon returning, I joined the WeBank and AIR Trustworthy Federated Learning Summer Camp, then went back to the institute. A few days later came the high-temperature holiday break. After that, it was National Day, followed by team-building events. This period felt like it flew by so fast; I barely have any memories of it. Later, I went to Hainan again to submit to IEEE ICPADS, and slipped off to Hainan again while writing this year-end summary.\nOther things that made me happy this year included buying a camera. Although I paid a bit of a \u0026ldquo;tuition fee\u0026rdquo; for it, feeling like a street wanderer was quite nice. I also visited many attractions in Beijing. In 2024, I bought an annual pass for Beijing parks. I hope to visit more places and earn back the ticket price QAQ.\nThis year, 2024, pressure has also arrived. Graduation-related matters have finally come to the forefront: IELTS scores, thesis, resume, various application documents, and information gathering are all tasks to be done. Online, I\u0026rsquo;m seeking a 25 Fall PhD position!!! If anyone is applying or wants to exchange ideas, feel free to find me via email, comments, or other methods. I\u0026rsquo;d be very happy to study and communicate together~!\nThis year, I hope my family stays healthy! :)\nAlso, wishing everyone a prosperous Year of the Dragon and all the best!\nSo, Happy New Year, and may everything go smoothly!\n","permalink":"https://blog.bj-yan.top/en/p/journey-annual-summary-2023/","summary":"\u003cp\u003eHi, long time no see. Today is the third day of the Lunar New Year, Happy New Year!\u003c/p\u003e\n\u003cp\u003eI\u0026rsquo;ve always wanted to document the fresh start of 2023. I thought about it a lot, but felt there might not be much worth recording. Still, since it\u0026rsquo;s been so long since I last wrote a blog post, I\u0026rsquo;ll just write something casually.\u003c/p\u003e\n\u003cp\u003eFirst, let\u0026rsquo;s casually talk about the year-end summaries from various apps. Bilibili is naturally my perfect-attendance champion! My watch time is almost 2000 hours. The year-end summaries from music apps are no longer very useful, since my music player has completely become YouTube and Bilibili. I can find whatever I want to listen to, there are no copyright issues, and I don\u0026rsquo;t pursue top-tier audio quality, so it\u0026rsquo;s just fine (maybe I should create a browser plugin to automatically track playback time?).\u003c/p\u003e","title":"Perhaps a 2023 Year-in-Review"},{"content":"Preface When reading federated learning papers, you often encounter a sentence: \u0026ldquo;Partition client data using a Dirichlet distribution and control the Non-IID degree with $\\alpha$.\u0026rdquo; It looks like a ready-made experimental switch, but if you only remember that \u0026ldquo;the smaller $\\alpha$ is, the more uneven the data is,\u0026rdquo; you easily miss critical details.\nIs this probability vector assigning different labels to a client, or assigning different clients to a label? Is $\\alpha$ the parameter for each coordinate, or the sum of parameters for all coordinates? Does the partitioning code inadvertently change the data volume per client? These questions all affect experimental results.\nThis article starts from the probability distribution itself and then provides a directly runnable data partitioning example. Here we discuss artificially constructed label heterogeneity, not a complete modeling of real-world federated data.\nWhy Simulate Non-IID? Suppose the data of client $k$ comes from distribution $P_k(X,Y)$. When distributions differ across clients, their local optimization objectives may also differ:\n$$ F_k(w)=\\mathbb E_{(X,Y)\\sim P_k}[\\ell(w;X,Y)], \\qquad F(w)=\\sum_{k=1}^{K}\\pi_k F_k(w). $$ Here, $\\pi_k\\ge0$, and $\\sum_k\\pi_k=1$; a common choice is weighting by the number of samples per client. Even if all clients start from the same model, after performing multiple steps of local training individually, the update directions may gradually diverge.\nBut Non-IID does not come in just one form:\nDifference An Intuitive Example Can Label-Only Partitioning Fully Simulate This? Label Distribution Difference $P_k(Y)$ Different ratios of cat vs. dog images on different devices Yes, this dimension can be constructed Conditional Feature Distribution Difference $P_k(X\\mid Y)$ For the same dog, different devices use different cameras to capture images No, cannot fully simulate Conditional Labeling Mechanism Difference $P_k(Y\\mid X)$ Different institutions adopt different labeling rules for similar samples No, cannot fully simulate Sample Size Difference $n_k$ Active users vs. low-frequency users have different data volumes May be simultaneously introduced by the partitioning process These descriptions are not entirely independent. For example, fixing the pools of samples for each category and changing label ratios can also alter the overall feature distribution of clients. In experiments, one should specify exactly what was constructed, rather than compressing all differences into a single \u0026ldquo;Non-IID\u0026rdquo; label.\nHsu et al. used the Dirichlet distribution to generate varying degrees of client label distributions to study the impact of such differences on FedAvg. It provides a useful experimental approach, but the meaning of its parameters must be read in conjunction with its specific construction. 1\nWhat is the Dirichlet Distribution? It Generates a Probability Vector A $d$-dimensional Dirichlet random vector satisfies\n$$ \\mathbf p=(p_1,\\ldots,p_d)\\sim\\operatorname{Dir}(\\alpha_1,\\ldots,\\alpha_d), \\qquad p_i\\ge0,\\quad\\sum_{i=1}^{d}p_i=1, $$ where each $\\alpha_i\u0026gt;0$. Its output is not a class label nor an integer sample count, but the proportion of each class or each client.\nLet $\\alpha_0=\\sum_i\\alpha_i$. Inside the simplex, its density is\n$$ f(\\mathbf p)=\\frac{\\Gamma(\\alpha_0)}{\\prod_{i=1}^{d}\\Gamma(\\alpha_i)} \\prod_{i=1}^{d}p_i^{\\alpha_i-1}. $$ Here, the density is defined relative to the $d-1$-dimensional coordinates, because the last component is determined by the preceding ones. In two dimensions, it degenerates to the Beta distribution; in three dimensions, each sample can be plotted as a point inside a triangle.\nEach point represents a set of proportions summing to 1. Points near a vertex indicate one component dominates, while points near the center indicate the three components are roughly equal.\nDefinitions and sampling methods can be found in Stanford Course Notes2 and NumPy Documentation3. In the figure, $\\alpha$ refers to the parameter for each coordinate.\nMean and Variance Are Two Different Things The mean, variance, and covariance between different coordinates of the Dirichlet distribution are\n$$ \\mathbb E[p_i]=\\frac{\\alpha_i}{\\alpha_0}, \\qquad \\operatorname{Var}(p_i)=\\frac{\\alpha_i(\\alpha_0-\\alpha_i)}{\\alpha_0^2(\\alpha_0+1)}, $$ $$ \\operatorname{Cov}(p_i,p_j)= -\\frac{\\alpha_i\\alpha_j}{\\alpha_0^2(\\alpha_0+1)},\\qquad i\\ne j. $$ Negative covariance is not surprising: since the sum of all proportions is fixed at 1, if one component grows larger, the others must yield space.\nIf symmetric parameters $\\alpha_1=\\cdots=\\alpha_d=\\alpha$ are used, then\n$$ \\mathbb E[p_i]=\\frac1d, \\qquad \\operatorname{Var}(p_i)=\\frac{d-1}{d^2(d\\alpha+1)}. $$ Therefore, reducing $\\alpha$ does not change the average proportion of each coordinate, but rather increases the fluctuation between a single sample and the average proportion. One sample might assign almost everything to the first component, while the next might assign almost everything to the third component.\nWhen $\\alpha\u0026lt;1$, the distribution is more biased toward the boundaries of the simplex, making it likely that some components are very small while others dominate. When $\\alpha=1$, the distribution is uniform over the simplex, but it does not output uniform proportions every time. When $\\alpha\u0026gt;1$, the distribution is more concentrated toward the center; when $\\alpha$ is very large, the proportions become closer to $1/d$. This is a trend at the distribution level and does not guarantee that every small $\\alpha$ sample is more skewed than every large $\\alpha$ sample. Even when the probability vector is very close to uniform, allocating a finite number of samples still introduces fluctuations; a large $\\alpha$ does not make every client\u0026rsquo;s empirical distribution exactly identical.\nSingle Coordinate Parameter vs. Total Concentration The parameters can also be written as $\\alpha_i=s m_i$, where $m_i\u0026gt;0$ and $\\sum_i m_i=1$, so\n$$ \\mathbf p\\sim\\operatorname{Dir}(s\\mathbf m), \\qquad \\mathbb E[p_i]=m_i,\\qquad\\alpha_0=s. $$ $\\mathbf m$ controls the center position, while $s$ controls the fluctuations around this center. A uniform center corresponds to $m_i=1/d$, in which case the parameter for each coordinate is $s/d$.\nTherefore, the parameters $\\operatorname{Dir}(\\alpha,\\ldots,\\alpha)$ and $\\operatorname{Dir}(\\alpha\\mathbf m)$ cannot be directly compared. The former has a total concentration of $d\\alpha$, while the latter has a total concentration of $\\alpha$. The paper by Hsu et al. uses the latter notation; when reproducing experiments, this distinction must be made explicit.\nIn Federated Learning, There Are Two Partitioning Directions Assume there are $K$ clients and $C$ classes, with the sample count for class $c$ being $N_c$.\nDirection 1: Generate Class Proportions for Each Client For each client $k$, sample a $C$-dimensional vector:\n$$ \\mathbf q_k\\sim\\operatorname{Dir}(s\\mathbf m). $$ This describes the class proportions the client is expected to hold. If the client requires a fixed number of $n_k$ samples, the integer counts for each class can then be determined.\nThe difficulty lies in the fact that the sum of demands across all clients may exceed the actual inventory of samples for a given class. If sampling independently from the class pool, one must clarify whether replacement is allowed; if not, supply-demand mismatches must be handled. Implementing balanced client sample sizes may also introduce additional constraints.\nDirection 2: Generate Client Proportions for Each Class For each class $c$, sample a $K$-dimensional vector:\n$$ \\mathbf r_c\\sim\\operatorname{Dir}(\\alpha,\\ldots,\\alpha), \\qquad (N_{1c},\\ldots,N_{Kc})\\mid\\mathbf r_c \\sim\\operatorname{Multinomial}(N_c,\\mathbf r_c). $$ In this way, all existing samples of each class can be distributed, and $\\sum_kN_{kc}=N_c$. The final data volume for each client is $n_k=\\sum_cN_{kc}$; for non-empty clients, their empirical label proportions are\n$$ \\widehat P_k(Y=c)=\\frac{N_{kc}}{n_k}. $$ Note that $r_{c,k}$ describes \u0026ldquo;the proportion of samples from class $c$ assigned to client $k$,\u0026rdquo; not \u0026ldquo;the proportion of class $c$ within client $k$.\u0026rdquo; Their normalization directions differ.\nWe adopt the second construction below. It preserves the total number of classes in the entire dataset but does not guarantee equal sample sizes per client, nor does it guarantee that every client has samples. When $\\alpha$ is small, a client might receive samples from multiple classes or none at all; one cannot simply copy the \u0026ldquo;one client, one class\u0026rdquo; intuition from the first construction.\nA Reproducible NumPy Implementation The function below accepts 1D integer labels and returns the original sample indices held by each client. It first shuffles the indices for each class, uses Dirichlet to generate proportions, and then determines integer sample counts via the multinomial distribution.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 import numpy as np def partition_dirichlet(labels, num_clients, alpha, seed=42): labels = np.asarray(labels) if labels.ndim != 1 or labels.size == 0: raise ValueError(\u0026#34;labels must be a non-empty 1D array\u0026#34;) if not np.issubdtype(labels.dtype, np.integer): raise ValueError(\u0026#34;labels must contain integer class IDs\u0026#34;) if isinstance(num_clients, bool) or not isinstance(num_clients, int) or num_clients \u0026lt; 2: raise ValueError(\u0026#34;num_clients must be an integer \u0026gt;= 2\u0026#34;) if not np.isscalar(alpha) or not np.isfinite(alpha) or alpha \u0026lt;= 0: raise ValueError(\u0026#34;alpha must be a finite positive number\u0026#34;) rng = np.random.default_rng(seed) parts = [[] for _ in range(num_clients)] for label in np.unique(labels): indices = rng.permutation(np.flatnonzero(labels == label)) proportions = rng.dirichlet(np.full(num_clients, alpha)) counts = rng.multinomial(indices.size, proportions) chunks = np.split(indices, np.cumsum(counts)[:-1]) for client, chunk in zip(parts, chunks): client.append(chunk) clients = [np.concatenate(part) for part in parts] for indices in clients: rng.shuffle(indices) return clients labels = np.repeat(np.arange(10), 500) clients = partition_dirichlet(labels, num_clients=20, alpha=0.5) # Every original sample appears exactly once. assigned = np.concatenate(clients) assert np.array_equal(np.sort(assigned), np.arange(labels.size)) # The same seed reproduces the same assignment in this environment. repeated = partition_dirichlet(labels, num_clients=20, alpha=0.5) assert all(np.array_equal(a, b) for a, b in zip(clients, repeated)) sizes = np.array([len(indices) for indices in clients]) print(\u0026#34;client sizes:\u0026#34;, sizes) print(\u0026#34;empty clients:\u0026#34;, np.count_nonzero(sizes == 0)) Here, we do not simply multiply the floating-point proportions by $N_c$ and floor the result, so there is no risk of \u0026ldquo;missing a few samples per class\u0026rdquo;; the sum of counts from the multinomial distribution exactly equals $N_c$. 4\nAn empty array is a valid partitioning result but cannot be directly fed into training pipelines that require non-empty datasets. If experiments disallow empty clients, one must redesign the allocation constraints or use resampling with a maximum count limit; these additional rules should be explicitly reported rather than silently discarding clients.\nThe label heatmap in the figure was generated using the same function: 20 clients, 10 classes, 500 samples per class, with a random seed of 42. Colors represent the class proportions within a non-empty client, and all panels use the same color scale.\nAfter assigning categories to clients, smaller alphas typically induce stronger label skew and may also alter the sample size per client. Empty clients, if any, are shown in gray.\nWhat Else to Check When Running Experiments? Do Not Just Save the Random Seed Recording the random seed is useful, but one should also save the final list of client indices, the dataset version, the sample order, and the NumPy version. Changing the data loading order or the sequence of random number calls can cause the same seed to produce completely different partitions.\nWhen comparing multiple algorithms under the same setting, try to reuse the same partition and report the mean and variance using multiple partition seeds. Do not mistake an accidentally easy partition for an algorithmic advantage.\nObserve Label Skew and Sample Size Skew Separately In addition to heatmaps, one should also statistics on the sample size, the number of non-empty classes, and label entropy for each client. For non-empty clients, label entropy can be written as\n$$ H_k=-\\sum_{c:\\widehat P_k(c)\u0026gt;0}\\widehat P_k(c)\\log\\widehat P_k(c). $$ Here we use the natural logarithm; for a uniform coverage of $C$ classes, it equals $\\log C$. Low entropy indicates more concentrated labels, but it does not fully describe differences between two clients: two clients may both have an entropy of 0 (each having only one class) while the classes themselves are completely different.\nAccuracy weighted by sample size and accuracy averaged across clients answer different questions. The former favors overall sample performance, while the latter focuses on the experience of an average client; when $n_k$ differences are large, the two metrics can diverge significantly.\nBe Careful About the Relationship Between Training, Validation, and Testing You can manually split only the training set and retain an independent, fixed global test set; if you need to evaluate per-client local performance, you must explicitly define how local test data is generated. Real-world multi-user data should also be split by user or entity to avoid highly correlated samples from the same individual crossing between training and test sets.\nSplitting training and validation data within pre-fixed client assignments, versus randomly splitting all samples first and then generating clients, may yield different evaluation targets. Both approaches must be clearly explained.\nComparative Experiments Must Report Full Configuration At minimum, record the split direction, definition of the parameter vector, $K$, $C$, sample sizes for each class, random seed, integer allocation method, handling of empty clients, minimum sample size constraints, as well as client sampling and aggregation weights. Reporting only \u0026ldquo;$\\alpha=0.5$\u0026rdquo; is insufficient for others to reproduce your experiment.\nConclusion The role of the Dirichlet distribution is to provide a probability vector whose fluctuation level can be tuned. It allows us to construct label heterogeneity more conveniently, but $\\alpha$ is not a universal difficulty scale that remains valid after removing dimensionality, allocation direction, and additional constraints.\nWhat truly deserves scrutiny is the final client data: which classes went where, how many samples each client holds, and whose performance the evaluation metrics are actually measuring. Clearly documenting these details is more important than simply placing a single Non-IID parameter in an experimental table.\nReferences and Further Reading Sources: 5.\nHsu, Qi, and Brown: Measuring the Effects of Non-Identical Data Distribution for Federated Visual Classification.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nStanford：Modelling Mixtures — the Dirichlet distribution。\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNumPy: Generator.dirichlet and Generator.multinomial.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNumPy: Generator.dirichlet and Generator.multinomial.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPrivacy Leakage in Deep Learning: Data remaining on clients does not imply that the transmitted updates lack sensitive information.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-dirichlet-distribution-in-federated-learning/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eWhen reading federated learning papers, you often encounter a sentence: \u0026ldquo;Partition client data using a Dirichlet distribution and control the Non-IID degree with $\\alpha$.\u0026rdquo; It looks like a ready-made experimental switch, but if you only remember that \u0026ldquo;the smaller $\\alpha$ is, the more uneven the data is,\u0026rdquo; you easily miss critical details.\u003c/p\u003e\n\u003cp\u003eIs this probability vector assigning different labels to a client, or assigning different clients to a label? Is $\\alpha$ the parameter for each coordinate, or the sum of parameters for all coordinates? Does the partitioning code inadvertently change the data volume per client? These questions all affect experimental results.\u003c/p\u003e","title":"The Dirichlet Distribution in Federated Learning: From Probability Vectors to Non-IID Data Partitioning"},{"content":"0x00 Preface The efficiency of distributed optimization requires measuring both iteration effectiveness and real-world elapsed time: how fast a single update completes, and how fast a certain accuracy is reached, are two distinct questions. This article combines the error bounds from Distributed Convergence Analysis 1 to analyze the time models for synchronous and asynchronous SGD.\nNotation: $m$ denotes the number of workers, $b$ the batch size per worker, $K$ the number of workers or batches received per update, $t$ the number of updates, and $T$ the wall-clock time. $L$ and $\\mu$ are the smoothness and strong convexity constants of the objective function, respectively; the exponential rate of service time is denoted by $\\lambda$ and should not be confused with the strong convexity constant.\n0x01 Service-time Model Let $X_i$ represent the time required for a worker to process a batch and return the gradient. To obtain a computable theoretical model, we assume service times are independent and identically distributed (i.i.d.) across workers and different tasks, and do not depend on the current sample or gradient noise. Server aggregation, broadcasting, and task cancellation overheads are not separately accounted for here; in real systems, these overheads must be measured and cannot be assumed to be included in each worker\u0026rsquo;s independent time.\nDistributions The probability density function is denoted by $p_X(x)$, and the cumulative distribution function by $F_X(x)=\\Pr(X\\le x)$.\nExponential distribution: $p_X(x)=\\lambda e^{-\\lambda x}$, $F_X(x)=1-e^{-\\lambda x}$, $x\\ge0$, $\\lambda\u0026gt;0$. Shifted exponential distribution: $X=\\Delta+E$, $E\\sim\\operatorname{Exp}(\\lambda)$, $\\Delta\\ge0$, therefore $\\mathbb E[X]=\\Delta+1/\\lambda$. Pareto distribution: When $x\\ge x_m\u0026gt;0$, $p_X(x)=\\alpha x_m^\\alpha/x^{\\alpha+1}$, $F_X(x)=1-(x_m/x)^\\alpha$; when $x\u0026lt;x_m$, both are zero. The mean of the Pareto distribution is $\\alpha x_m/(\\alpha-1)$ when $\\alpha\u0026gt;1$, and diverges when $0\u0026lt;\\alpha\\le1$; the variance is $\\alpha x_m^2/((\\alpha-1)^2(\\alpha-2))$ when $\\alpha\u0026gt;2$. When $1\u0026lt;\\alpha\\le2$, the variance is infinite; when the mean is infinite, the variance is not treated as a quantity under a finite mean model.\nThe standard Gaussian distribution allows negative values and cannot directly serve as a strict runtime model. If used to approximate measured values, one must specify the probability of negative values or use a truncated distribution. The exact harmonic number formulas below apply only to exponential or specifically noted shifted exponential models, not to arbitrary service distributions.\nOrder Statistics Sort the $m$ runtimes as $X_{1:m}\\le\\cdots\\le X_{m:m}$. $X_{K:m}$ is the K-th smallest value among all m times, not the maximum time of a pre-specified set of K workers.\nLet $H_n=\\sum_{r=1}^n1/r$, $H_0=0$. For $m$ independent exponential variables, the expected intervals between adjacent order statistics are $1/(m\\lambda),1/((m-1)\\lambda),\\ldots$, so\n$$ \\mathbb E[X_{K:m}]=\\frac{H_m-H_{m-K}}{\\lambda}. $$ If a constant offset $\\Delta$ is added to each time, the corresponding result also increases by $\\Delta$.\n0x02 Synchronous SGD Time per Update Synchronous updates wait for all workers, therefore\n$$ S_t=\\max_{1\\le i\\le m}X_{i,t},\\qquad s_m:=\\mathbb E[S_t]=\\Delta+\\frac{H_m}{\\lambda}. $$ $H_m=\\log m+\\gamma_{\\rm E}+O(1/m)$, where $\\gamma_{\\rm E}$ is the Euler constant. $\\log m$ is the dominant term only for large m; when $m=1$, one must use $H_1=1$.\nThe total time to complete a fixed number t of updates is $W_t=\\sum_{j=0}^{t-1}S_j$, hence\n$$ \\mathbb E[W_t]=ts_m. $$ If each round\u0026rsquo;s time is i.i.d. with finite mean, the number of updates completed within a fixed time $N(T)$ satisfies $N(T)/T\\to1/s_m$ (almost surely). The long-term batch throughput is $m/s_m$. This is not the exact expectation of an arbitrary throughput random variable within a finite time.\nExpected Time to Complete Enough Updates Assume the target is $L$-smooth and $\\mu$-strongly convex, with gradients satisfying the unbiased and independent noise conditions. With a fixed step size $0\u0026lt;\\eta\\le1/L$, let\n$$ \\mathcal E_t=\\mathbb E[f(w_t)]-f^*,\\quad q=1-\\eta\\mu,\\quad C_m=\\frac{\\eta L\\sigma^2}{2\\mu bm}. $$ Then $\\mathcal E_t\\le C_m+q^t(\\mathcal E_0-C_m)$. If $0\u0026lt;q\u0026lt;1$ and $\\mathcal E_0\u0026gt;\\epsilon\u0026gt;C_m$, define\n$$ t_\\epsilon=\\left\\lceil \\frac{\\log((\\mathcal E_0-C_m)/(\\epsilon-C_m))}{-\\log q} \\right\\rceil. $$ After completing these t updates, the expected function value error does not exceed $\\epsilon$, with an expected elapsed time of $s_m t_\\epsilon$. This is not the expected hitting time for a specific random training trajectory to first reach the error threshold.\nIf $\\epsilon\\le C_m$, the current bound cannot guarantee reaching the target. If the initial error already satisfies the target, one may stop; when using $q=0$, employ single-step recursion without computing logarithms.\nChoosing a Step Size for a Target Accuracy Let $\\sigma^2\u0026gt;0$ and $\\mathcal E_0\u0026gt;\\epsilon\u0026gt;0$. One may choose\n$$ \\eta=\\min\\left\\{\\frac1L,\\frac{\\mu bm\\epsilon}{L\\sigma^2}\\right\\}, $$ to make $C_m\\le\\epsilon/2$. Utilizing $q^t\\le e^{-\\eta\\mu t}$, take\n$$ t\\ge\\left\\lceil\\frac1{\\eta\\mu}\\log\\frac{2\\mathcal E_0}{\\epsilon}\\right\\rceil $$ sufficient to reach the target. Thus, we can provide an upper bound on runtime retaining the deterministic term, the noise term, and the logarithmic term:\n$$ \\mathbb E[W_t] =O\\left[ \\left(\\Delta+\\frac{H_m}{\\lambda}\\right) \\left(1+\\max\\left\\{\\frac L\\mu,\\frac{L\\sigma^2}{\\mu^2bm\\epsilon}\\right\\} \\log\\frac{2\\mathcal E_0}{\\epsilon}\\right) \\right]. $$ Approximate iterative scaling of $1/(m\\epsilon)$ occurs only within the corresponding range where the noise term dominates, other fixed constants are ignored, and necessary logarithmic terms are retained. Once m increases to the point where the step-size upper bound becomes active, the number of iterations no longer decreases indefinitely according to $1/m$; the waiting cost per round still increases. When $\\sigma^2=0$, directly adopt the complexity of deterministic strongly convex GD.\nError at a Fixed Wall-clock Time $N(T)$ is a random variable; one cannot replace it with $T/s_m$ and still write the formula as a strict upper bound. In models where the runtime process and gradient sampling are independent, one should first condition on $N(T)$:\n$$ \\mathbb E[f(w_{N(T)})]-f^* \\le C_m+(\\mathcal E_0-C_m)\\mathbb E[q^{N(T)}]. $$ Even if these independence assumptions hold, one cannot directly substitute $\\mathbb E[q^{N(T)}]$ with $q^{\\mathbb E[N(T)]}$. For $0\u0026lt;q\u0026lt;1$, the latter is a lower bound of the former by Jensen\u0026rsquo;s inequality; when $\\mathcal E_0\u0026gt;C_m$, this substitution would underestimate the current error upper bound.\nFor example, when $m=1,\\Delta=0$, $N(T)$ is a Poisson variable with parameter $\\lambda T$, hence $\\mathbb E[q^{N(T)}]=\\exp(\\lambda T(q-1))$ generally does not equal $q^{\\lambda T}$. $N(T)\\approx T/s_m$ can be used to illustrate trends but must be labeled as an approximation.\n0x03 K-Sync and K-Batch-Sync SGD K-Sync Each round starts from the same parameter, receiving results from the fastest K distinct workers and canceling the remaining tasks. Therefore, for shifted exponential latency,\n$$ s_K:=\\mathbb E[S_{K\\text{-sync}}] =\\Delta+\\frac{H_m-H_{m-K}}{\\lambda},\\qquad1\\le K\\le m. $$ when $K=1$ it becomes $\\Delta+1/(m\\lambda)$, and when $K=m$ it reverts to synchronous results. The logarithmic approximation $\\lambda^{-1}\\log(m/(m-K))$ does not apply to $K=m$ and should also be avoided near boundaries in place of exact harmonic numbers.\nUnder models where the selected batch remains conditionally unbiased and noise-independent, the error bound uses $C_K=\\eta L\\sigma^2/(2\\mu bK)$. Given the same target error, compare $s_Kt_{\\epsilon,K}$ rather than only comparing $s_K$. Choosing a smaller K can reduce waiting time but increases the noise term in the bound. If speed correlates with samples or local data distributions, selection bias must be addressed first.\nK-Batch-Sync All workers compute on the current parameter; once a worker finishes, it continues computing the next independent batch until a total of K batches are received, after which an update occurs and unfinished old tasks are canceled.\nUnder shifted exponential service times, immediate restart, and zero cancellation overhead, m independent Poisson completion processes superimpose; each round waits for K completion events, hence\n$$ \\mathbb E[S_{K\\text{-batch-sync}}]=\\frac K{m\\lambda}. $$ Here K is the number of batches and need not be less than m. Under shifted or general service times, the completion process within a round is typically not a Poisson process, so this formula cannot be used directly.\n0x04 Asynchronous SGD and Renewal Theory In single-gradient asynchronous SGD, workers upload immediately upon completion, read the current parameter, and start the next batch; the server updates once for every gradient received. Server queuing and model transmission bottlenecks are not considered here.\nLet worker i\u0026rsquo;s consecutive service times $X_{i,1},X_{i,2},\\ldots$ be independent and identically distributed, strictly positive, and $0\u0026lt;\\mathbb E[X]\u0026lt;\\infty$. Let\n$$ J_{i,n}=\\sum_{r=1}^nX_{i,r},\\qquad A_i(T)=\\max\\{n\\ge0:J_{i,n}\\le T\\},\\qquad J_{i,0}=0. $$ This defines the renewal process. The renewal count limit derived from the strong law of large numbers is $A_i(T)/T\\to1/\\mathbb E[X]$ (almost surely); the basic renewal theorem provides the expected version $\\mathbb E[A_i(T)]/T\\to1/\\mathbb E[X]$. These two should be distinguished.\nThe long-term total update rate for all workers is $m/\\mathbb E[X]$, so the long-term average time per update is\n$$ \\bar s_{\\rm async}=\\frac{\\mathbb E[X]}m. $$ It does not guarantee that the expected waiting time for each update is identical from initialization. For example, if all service times are fixed at 1, the first completion event still waits 1; afterward, simultaneous arrival events may occur.\nComparing Sync and Async Under non-shifted exponential times, the waiting time for single-gradient asynchronous updates is $\\operatorname{Exp}(m\\lambda)$, hence the mean is $1/(m\\lambda)$. The synchronous mean per round is $H_m/\\lambda$, with a ratio of $mH_m$.\nUnder the shifted exponential model, comparing the long-term asynchronous average with the synchronous single-round mean yields\n$$ \\frac{s_m}{\\bar s_{\\rm async}} =m\\frac{\\lambda\\Delta+H_m}{\\lambda\\Delta+1}. $$ Asynchronous single updates are faster, but synchronous updates use m batches per round while asynchronous uses one batch per update, and asynchronous suffers from stale gradients. This time ratio is not the training speedup to achieve the same accuracy. Asynchronous error bounds also depend on the relative staleness and conditional noise assumptions from Part 4 1.\n0x05 K-Async and K-Batch-Async SGD K-Async K-Async receives results from K distinct workers each time. Workers that have already returned results in the current round wait for the server\u0026rsquo;s update before reading new parameters; those not yet finished continue their old computations, so the remaining service time at the start of the next round typically differs from a fresh service time.\nUnder independent, homogeneous, non-shifted exponential service times, the memoryless property ensures the remaining time is still an independent exponential variable, hence the mean per round is\n$$ \\mathbb E[S_{K\\text{-async}}]=\\frac{H_m-H_{m-K}}\\lambda,\\qquad1\\le K\\le m. $$ The general distribution cannot use this equality. If X satisfies the new-longer-than-used condition, meaning that for all non-negative a, u (when the conditional event probability is non-zero),\n$$ \\Pr(X\u0026gt;a+u\\mid X\u0026gt;a)\\le\\Pr(X\u0026gt;u), $$ then the remaining time of the running task is stochastically less than or equal to the fresh time, yielding $\\mathbb E[S_{K\\text{-async}}]\\le\\mathbb E[X_{K:m}]$. For example, the shifted exponential distribution satisfies this condition, with the right-hand side being $\\Delta+(H_m-H_{m-K})/\\lambda$, but generally there is no equality. Any other distribution still requires computing its own order statistics rather than using the exponential harmonic number formula.\nK-Batch-Async This variant updates after the server receives every K batches; workers read the current parameters and continue working upon completing each batch, without requiring contributors to be K distinct workers and without canceling in-progress batches.\nUnder the ideal unbiased exponential model, the mean time to wait for K total completion events is $K/(m\\lambda)$. For general finite-mean service times, renewal theory provides the long-run average time per server update, $K\\mathbb E[X]/m$, rather than an exact expectation equation for any finite number of rounds.\nK-Async and K-batch-async have different waiting and lagging mechanisms. When comparing them, one should simultaneously verify their respective conditionally unbiasedness, noise correlation, and lag bounds, and then compare the time required to achieve the same target accuracy.\n0x06 Reference This paper treats the expected completion time for a fixed number of updates, the expected error within a fixed time, and the long-term throughput separately, avoiding treating asymptotic or mean approximations as strict guarantees for finite time.\nSources: 2 3 4 5.\nDistributed Convergence Analysis \u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDutta, S., Joshi, G., Ghosh, S., Dube, P. and Nagpurkar, P., Slow and Stale Gradients Can Win the Race: Error-Runtime Trade-offs in Distributed SGD, AISTATS 2018, Section 4 and Supplement, Section 7.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDutta, S., Joshi, G., Ghosh, S., Dube, P. and Nagpurkar, P., Slow and Stale Gradients Can Win the Race: Error-Runtime Trade-offs in Distributed SGD, AISTATS 2018, Section 4 and Supplement, Section 7.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJoshi, G., Optimization Algorithms for Distributed Machine Learning, Springer.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGallager, R. G., Discrete Stochastic Processes, Chapter 4: Renewal Processes.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-efficiency-analysis-of-distributed-machine-learning/","summary":"\u003ch2 id=\"0x00-preface\"\u003e0x00 Preface\u003c/h2\u003e\n\u003cp\u003eThe efficiency of distributed optimization requires measuring both iteration effectiveness and real-world elapsed time: how fast a single update completes, and how fast a certain accuracy is reached, are two distinct questions. This article combines the error bounds from Distributed Convergence Analysis \u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e to analyze the time models for synchronous and asynchronous SGD.\u003c/p\u003e\n\u003cp\u003eNotation: $m$ denotes the number of workers, $b$ the batch size per worker, $K$ the number of workers or batches received per update, $t$ the number of updates, and $T$ the wall-clock time. $L$ and $\\mu$ are the smoothness and strong convexity constants of the objective function, respectively; the exponential rate of service time is denoted by $\\lambda$ and should not be confused with the strong convexity constant.\u003c/p\u003e","title":"Efficiency Analysis in Distributed Machine Learning"},{"content":"0x00 Preface Part 31 analyzed mini-batch SGD, Part 42 discussed synchronous and asynchronous distributed updates. This article turns to federated learning: clients not only hold different data but also perform multiple consecutive updates between a single communication round. Therefore, local stochastic gradients being unbiased for the local objective does not imply that the aggregated direction is unbiased for the current global gradient.\nStarting from FedAvg\u0026rsquo;s update rule, this article provides a simplified proof that can be verified step-by-step: first analyzing the model under full participation, equal-weight clients, and identical local steps, then deriving extensions for uniform partial participation and discussing inconsistent aggregation weights. The non-convex part guarantees the average gradient norm decreases; the strongly convex part additionally employs the PL inequality. These two conclusions must not be mixed.\nThe original FedAvg algorithm is described in McMahan et al., 20173. The constant bounds below are conservative bounds derived by this paper under explicit assumptions; they are not copied from any theorem in another paper, nor do they pursue optimal complexity.\n0x01 Objective and FedAvg Updates Global Objective Suppose there are $N$ clients, and the local objective for the $i$-th client is\n$$ f_i(w)=\\mathbb E_{\\xi\\sim\\mathcal D_i}[\\ell(w;\\xi)],\\qquad f(w)=\\sum_{i=1}^N p_i f_i(w),\\qquad p_i\u0026gt;0,\\quad\\sum_{i=1}^Np_i=1. $$ If one wishes for each sample to have equal weight, typically set $p_i=n_i/\\sum_j n_j$; if one wishes for each client to have equal weight, then set $p_i=1/N$. These are two distinct optimization objectives, not interchangeable implementation details. Data not being uploaded does not automatically imply differential privacy guarantees; this article only discusses optimization error.\nThe main theorem later fixes the use of $p_i=1/N$. Weighted objectives and partial participation are discussed later and are not included by default in the main theorem.\nLocal Steps and Communication Rounds $t$ denotes communication rounds, and $k$ denotes local update steps within a round. All clients start each round from the same server model and execute $E\\ge1$ local SGD steps:\n$$ \\begin{aligned} w_{i,t}^{(0)}\u0026amp;=w_t,\\\\ w_{i,t}^{(k+1)}\u0026amp;=w_{i,t}^{(k)}-\\eta g_{i,t}^{(k)},\\qquad k=0,\\ldots,E-1,\\\\ w_{t+1}\u0026amp;=\\frac1N\\sum_{i=1}^Nw_{i,t}^{(E)}. \\end{aligned} $$ $\\eta$ is the fixed local learning rate; the server directly averages the models without multiplying by an additional server learning rate. Define the effective step size per round as $\\alpha=\\eta E$, and the aggregation direction as\n$$ v_t=\\frac1{NE}\\sum_{i=1}^N\\sum_{k=0}^{E-1}g_{i,t}^{(k)}, \\qquad w_{t+1}=w_t-\\alpha v_t. $$ Here, $E$ is the number of updates, not the number of local epochs. If clients have different data volumes, the same number of epochs often corresponds to different numbers of updates and cannot be directly substituted into this model. Over a total of $T$ rounds, each client performs $ET$ local updates; the entire system computes a total of $NET$ batch gradients.\n0x02 Assumptions and Error Decomposition Smoothness and Lower Bound Each $f_i$ is a $L$-smooth function on Euclidean space, so the average objective $f$ is also $L$-smooth. Assume $f(w)\\ge f_{\\inf}\u0026gt;-\\infty$, define $\\Delta_0=f(w_0)-f_{\\inf}$, and treat $w_0$ as a deterministic initial value. Non-convex analysis does not require $f_{\\inf}$ to be attained by any parameter.\nConditional Gradient Noise Let $\\mathcal F_t$ contain the history before the start of round $t$; $\\mathcal H_{t,k}$ further includes all clients\u0026rsquo; parameters and history before the sampling of step $k$ within that round. Write\n$$ g_{i,t}^{(k)}=\\nabla f_i(w_{i,t}^{(k)})+\\varepsilon_{i,t}^{(k)}. $$ $$ \\mathbb E[\\varepsilon_{i,t}^{(k)}\\mid\\mathcal H_{t,k}]=0,\\qquad \\mathbb E[\\|\\varepsilon_{i,t}^{(k)}\\|^2\\mid\\mathcal H_{t,k}] \\le\\frac{\\sigma^2}{b}. $$ $b$ is the batch size per step, and $\\sigma^2$ is a uniform upper bound on the variance of single-sample noise. For the same step across different clients, the noise is also required to be conditionally uncorrelated. No unconditional independence is required between different steps of the same client; conditional unbiasedness ensures that noise increments form a martingale difference sequence.\nThe reduction of $1/b$ within a batch requires independent sampling or corresponding covariance conditions. If directly traversing the same random permutation, selecting samples based on loss, or if the sampling process is affected by dropout mechanisms, the above conditions must be re-verified; one cannot assume the conditions hold simply because \u0026ldquo;SGD was used\u0026rdquo;.\nBounded Gradient Heterogeneity Adopt the following unified heterogeneity upper bound:\n$$ \\frac1N\\sum_{i=1}^N\\|\\nabla f_i(w)-\\nabla f(w)\\|^2\\le\\zeta^2, \\qquad \\forall w. $$ $\\sigma^2$ describes random noise within a single client, while $\\zeta^2$ describes target differences between clients. Increasing the batch size reduces the former but does not automatically eliminate the latter. From $\\sum_i(\\nabla f_i-\\nabla f)=0$, we have\n$$ \\frac1N\\sum_i\\|\\nabla f_i(w)\\|^2 =\\|\\nabla f(w)\\|^2+ \\frac1N\\sum_i\\|\\nabla f_i(w)-\\nabla f(w)\\|^2 \\le\\|\\nabla f(w)\\|^2+\\zeta^2. $$ This is a strong assumption, not automatically implied by the description \u0026ldquo;non-IID data\u0026rdquo;. It requires a uniform bound on differences across all parameters, not just near the optimal point. If it holds only in a specific region, one must also prove that all relevant iterations remain within that region. This article does not take \u0026ldquo;globally uniformly bounded gradients\u0026rdquo; as an additional assumption.\nNoise and Local Drift Let $h_t=\\nabla f(w_t)$, and decompose the aggregation direction precisely into\n$$ \\begin{aligned} v_t\u0026amp;=h_t+r_t+e_t,\\\\ r_t\u0026amp;=\\frac1{NE}\\sum_{i,k} \\left[\\nabla f_i(w_{i,t}^{(k)})-\\nabla f_i(w_t)\\right],\\\\ e_t\u0026amp;=\\frac1{NE}\\sum_{i,k}\\varepsilon_{i,t}^{(k)}. \\end{aligned} $$ $r_t$ is the drift caused by the local path, and $e_t$ is the average random noise. Conditional unbiasedness and mutual uncorrelatedness yield\n$$ \\mathbb E[e_t\\mid\\mathcal F_t]=0,\\qquad \\mathbb E[\\|e_t\\|^2\\mid\\mathcal F_t]\\le \\frac{\\sigma^2}{bNE}. $$ The cross terms of asynchronous noise vanish via iterated conditional expectations. However, $r_t$ depends on local parameters, which in turn depend on previous noise, so one generally cannot assert $\\mathbb E\\langle r_t,e_t\\rangle=0$. The subsequent proof retains this correlation and controls it via norm inequalities.\n0x03 Bounding Local Drift This section provides the key estimates required for the theorems that follow. Let $q=E-1$, and define the average displacement under the history at the start of the round:\n$$ D_{t,k}=\\frac1N\\sum_i \\mathbb E[\\|w_{i,t}^{(k)}-w_t\\|^2\\mid\\mathcal F_t], \\qquad M_t=\\max_{0\\le k\\le q}D_{t,k}. $$ Here we only need to control the parameters of $k=0,\\ldots,E-1$, as these are the points where gradients are computed. When $E=1$, we have $q=0$, and thus $M_t=r_t=0$.\nExpand the local updates to separate the gradient mean from the noise. Using $|a+c|^2\\le2|a|^2+2|c|^2$, $|\\sum_{s\u0026lt;k}a_s|^2\\le k\\sum_{s\u0026lt;k}|a_s|^2$, and the martingale difference property of the noise, we obtain\n$$ D_{t,k}\\le \\frac{2\\eta^2k}{N}\\sum_i\\sum_{s\u0026lt;k} \\mathbb E[\\|\\nabla f_i(w_{i,t}^{(s)})\\|^2\\mid\\mathcal F_t] +\\frac{2\\eta^2k\\sigma^2}{b}. $$ Then apply smoothness and the heterogeneity bound from the previous section:\n$$ \\frac1N\\sum_i \\mathbb E[\\|\\nabla f_i(w_{i,t}^{(s)})\\|^2\\mid\\mathcal F_t] \\le2(\\|h_t\\|^2+\\zeta^2)+2L^2D_{t,s}. $$ $$ D_{t,k}\\le4\\eta^2k^2(\\|h_t\\|^2+\\zeta^2) +4\\eta^2kL^2\\sum_{s\u0026lt;k}D_{t,s} +\\frac{2\\eta^2k\\sigma^2}{b}. $$ Taking the maximum over $0\\le k\\le q$ yields\n$$ M_t\\le4\\eta^2q^2(\\|h_t\\|^2+\\zeta^2) +4\\eta^2L^2q^2M_t+\\frac{2\\eta^2q\\sigma^2}{b}. $$ From now on, we uniformly adopt the step-size condition $0\u0026lt;\\eta LE\\le1/8$. Thus $4\\eta^2L^2q^2\\le1/16$; rearranging terms and loosening the constants gives\n$$ M_t\\le8\\eta^2q^2(\\|h_t\\|^2+\\zeta^2) +\\frac{4\\eta^2q\\sigma^2}{b}. $$ By Jensen\u0026rsquo;s inequality and the smoothness of the local objectives:\n$$ \\begin{aligned} \\mathbb E[\\|r_t\\|^2\\mid\\mathcal F_t] \u0026amp;\\le\\frac{L^2}{E}\\sum_{k=0}^{E-1}D_{t,k}\\\\ \u0026amp;\\le8L^2\\eta^2q^2(\\|h_t\\|^2+\\zeta^2) +\\frac{4L^2\\eta^2q\\sigma^2}{b}. \\end{aligned} $$ This also explains why \u0026ldquo;treating the $NEb$ samples in one round as a single large batch\u0026rdquo; is insufficient: while the noise average does enjoy a $1/(NE)$ reduction, the gradients are computed at different local parameters, so one must also control $r_t$.\n0x04 A Non-convex Convergence Bound One-round Descent Using the smoothness of $f$ and $w_{t+1}=w_t-\\alpha v_t$,\n$$ f(w_{t+1})\\le f(w_t)-\\alpha\\langle h_t,v_t\\rangle +\\frac{L\\alpha^2}{2}\\|v_t\\|^2. $$ First take the expectation conditioned on $\\mathcal F_t$. $h_t$ is measurable, $\\mathbb E[e_t\\mid\\mathcal F_t]=0$, so the inner product of $h_t$ with the noise vanishes. For the drift term, use\n$$ |\\langle h_t,r_t\\rangle|\\le\\frac14\\|h_t\\|^2+\\|r_t\\|^2, \\qquad \\|h_t+r_t+e_t\\|^2\\le4\\|h_t\\|^2+4\\|r_t\\|^2+2\\|e_t\\|^2. $$ Here we do not assume that $r_t$ and $e_t$ are independent. Denoting $R_t=\\mathbb E[|r_t|^2\\mid\\mathcal F_t]$, we have\n$$ \\mathbb E[f(w_{t+1})\\mid\\mathcal F_t] \\le f(w_t) -\\alpha\\left(\\frac34-2L\\alpha\\right)\\|h_t\\|^2 +\\alpha(1+2L\\alpha)R_t +\\frac{L\\alpha^2\\sigma^2}{bNE}. $$ Since $L\\alpha=L\\eta E\\le1/8$, the descent coefficient of the first term is at least $1/2$, and $1+2L\\alpha\\le5/4$. Substituting the drift bound, the new coefficient for the gradient term is at most\n$$ 10L^2\\eta^2q^2\\le\\frac{10}{64}\u0026lt;\\frac14. $$ Therefore, we can conservatively retain a descent amount of $\\alpha/4$. Define\n$$ A=\\frac{L\\eta\\sigma^2}{bN} +\\frac{5L^2\\eta^2(E-1)\\sigma^2}{b} +10L^2\\eta^2(E-1)^2\\zeta^2. $$ This yields the core recursion of the entire proof:\n$$ \\mathbb E[f(w_{t+1})\\mid\\mathcal F_t] \\le f(w_t)-\\frac\\alpha4\\|\\nabla f(w_t)\\|^2+\\alpha A. $$ Stationarity Guarantee defination\nUnder the aforementioned assumptions, equal-weight aggregation with full participation, and $0\u003c\\eta\\le1/(8LE)$, running $T\\ge1$ communication rounds gives $$ \\begin{aligned} \\frac1T\\sum_{t=0}^{T-1}\\mathbb E\\|\\nabla f(w_t)\\|^2 \u0026amp;\\le\\frac{4\\Delta_0}{\\eta ET}+4A\\\\ \u0026amp;=\\frac{4\\Delta_0}{\\eta ET} +\\frac{4L\\eta\\sigma^2}{bN} +\\frac{20L^2\\eta^2(E-1)\\sigma^2}{b} +40L^2\\eta^2(E-1)^2\\zeta^2. \\end{aligned} $$ The proof only requires taking the total expectation of the core recursion and summing over $t=0,\\ldots,T-1$: the function value terms cancel out, and then we apply $\\mathbb E[f(w_T)]\\ge f_{\\inf}$. If we independently and uniformly select one $w_R$ from these $T$ initial parameters of the rounds, the left-hand side equals exactly $\\mathbb E|\\nabla f(w_R)|^2$.\nThis is an average stationary point bound, not a guarantee for the last round\u0026rsquo;s $w_T$, nor a guarantee of reaching the global optimum. A small gradient does not imply high test accuracy.\nReading the Four Terms Term Meaning Implication visible from this bound $4\\Delta_0/(\\eta ET)$ Optimization term for finite rounds With other conditions fixed, more rounds reduce this term $4L\\eta\\sigma^2/(bN)$ Gradient noise after aggregation More independent clients or larger batches can reduce this term $20L^2\\eta^2(E-1)\\sigma^2/b$ Drift accumulated by noise along local paths Cannot be eliminated solely by the $1/N$ reduction in aggregation $40L^2\\eta^2(E-1)^2\\zeta^2$ Drift caused by heterogeneous objectives Performing more local steps may increase this term Under a fixed step size, the last three terms do not automatically vanish by increasing $T$. They are residuals in the upper bound, not precise lower bounds on the algorithm\u0026rsquo;s actual error. Even if $\\zeta=0$, random noise may still produce drift via different local paths; even if $\\sigma=0$, heterogeneity may still cause bias.\nIf we fix $N,E,b$ and choose $\\eta=c/\\sqrt T$ for a run of length $T$, where $0\u0026lt;c\\le1/(8LE)$, this bound yields $O(T^{-1/2})+O(T^{-1})$. Here, a fixed step size is still used within each run; this is not a proof for online-changing $\\eta_t$. When $E$ or $N$ grows with $T$, one must re-substitute all terms and cannot continue to hide them in constants.\nWhen $E=1$, the drift terms vanish completely, and the algorithm recovers full-participation synchronous mini-batch SGD. The constants in the above general bound are relatively loose; directly using the unbiased gradient proof from Part 3 1 yields better constants. If the squared norm of the objective gradient is $\\epsilon$, this bound can only provide a guarantee when $\\epsilon\u0026gt;4A$ by increasing $T$.\n0x05 Strongly Convex Objectives If we further assume $f$ is $\\mu$-strongly convex and an optimal solution $w^\\ast$ exists, let $f^\\ast=f(w^\\ast)$ and $\\mathcal E_t=\\mathbb E[f(w_t)]-f^\\ast$. Applying the PL inequality\n$$ \\|\\nabla f(w)\\|^2\\ge2\\mu(f(w)-f^\\ast), $$ The core recurrence becomes\n$$ \\mathcal E_{t+1}\\le \\left(1-\\frac{\\alpha\\mu}{2}\\right)\\mathcal E_t+\\alpha A. $$ Since $\\mu\\le L$, the step-size condition guarantees $\\rho=1-\\alpha\\mu/2\\in(0,1)$. Expanding the geometric series:\ndefination\n$$ \\mathcal E_T\\le\\rho^T\\mathcal E_0 +\\frac{2A}{\\mu}(1-\\rho^T). $$ With fixed local step sizes, the transient term contracts geometrically, but this bound typically only guarantees $\\limsup_T\\mathcal E_T\\le2A/\\mu$. Only when the residual vanishes does this recurrence yield geometric convergence to the exact optimal value.\nIf $0\u0026lt;2A/\\mu\u0026lt;\\epsilon\u0026lt;\\mathcal E_0$, the sufficient number of rounds is\n$$ T\\ge\\left\\lceil \\frac{\\log\\left((\\mathcal E_0-2A/\\mu)/(\\epsilon-2A/\\mu)\\right)}{-\\log\\rho} \\right\\rceil. $$ When $A=0$, directly use $\\rho^T\\mathcal E_0$; when the objective falls below the residual upper bound, the above round formula cannot provide guarantees. For a more detailed strongly convex analysis with decaying step sizes, see Li et al., ICLR 20204, but note that the assumptions, time indexing, and aggregation scheme in that paper require item-by-item alignment; one cannot simply replace the theorem here with a single $O(1/T)$ conclusion from it.\n0x06 A Deterministic Drift Example The residual mentioned earlier is merely an upper bound. Below, we use a noiseless example to show that local drift can genuinely exist, rather than being an artifact of proof techniques.\nConsider two equally weighted clients, $0\u0026lt;a\u0026lt;1/\\sqrt2$:\n$$ f_1(w)=\\frac12w^2+a(\\sin w+\\cos w),\\qquad f_2(w)=\\frac12w^2-a(\\sin w+\\cos w). $$ $$ f(w)=\\frac12w^2,\\qquad w^\\ast=0,\\qquad \\nabla f_{1,2}(w)=w\\pm a(\\cos w-\\sin w). $$ Both local objectives are $L$-smooth and strongly convex; we can take $L=1+\\sqrt2a$ and local strong convexity parameter $1-\\sqrt2a$. The heterogeneity bound holds over the entire space:\n$$ \\frac12\\sum_{i=1}^2|\\nabla f_i(w)-\\nabla f(w)|^2 =a^2(\\cos w-\\sin w)^2\\le2a^2. $$ Starting from the global optimum $w_t=0$, each client performs two steps of deterministic GD. The first step reaches $-\\eta a$ and $\\eta a$ respectively; averaging the models after the second step yields exactly\n$$ w_{t+1}=-\\eta a\\sin(\\eta a). $$ For sufficiently small positive step sizes, this is not zero: FedAvg can even move away from the global optimum. Conversely, when $E=1$, the two local gradients at $w=0$ completely cancel out.\nIf $0\u0026lt;\\eta\\le1/L$, each local GD mapping is a contraction; the average of the two-step mappings remains a contraction. Thus, this fixed-step-size FedAvg mapping has a unique attracting fixed point, and since it does not map zero to zero, the fixed point is not equal to the global optimum. This example satisfies the heterogeneity assumptions of the main theorem; additionally, taking $\\eta\\le1/(16L)$ also satisfies the main theorem\u0026rsquo;s step-size constraints for $E=2$.\nThe following Python code verifies the one-round formula and the fixed point; it runs directly without third-party libraries:\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 from math import cos, sin, sqrt a, eta = 0.25, 0.04 assert eta * 2 * (1 + sqrt(2) * a) \u0026lt;= 1 / 8 def fedavg_round(w, steps=2): local = [] for sign in (1, -1): x = w for _ in range(steps): x -= eta * (x + sign * a * (cos(x) - sin(x))) local.append(x) return sum(local) / 2 assert abs(fedavg_round(0) + eta * a * sin(eta * a)) \u0026lt; 1e-12 assert abs(fedavg_round(0, steps=1)) \u0026lt; 1e-12 w = 0.0 for _ in range(2000): w = fedavg_round(w) assert abs(w) \u0026gt; 1e-6 assert abs(fedavg_round(w) - w) \u0026lt; 1e-12 print(f\u0026#34;FedAvg fixed point: {w:.8f}; global optimum: 0\u0026#34;) This is a numerical check for the analytical example, not a substitute for a general convergence proof via experiments.\n0x07 Partial Participation and Aggregation Weights Client Sampling Adds Another Variance Term Now consider only $E=1$, still optimizing the equally weighted objective. Each round, we select $m$ clients uniformly and without replacement from $N\u0026gt;1$ clients. Selection occurs before batch sampling and is independent of the noise in the new samples for that round. If the aggregation direction is $\\hat g_t=m^{-1}\\sum_{i\\in S_t}g_i(w_t)$, then\n$$ \\mathbb E[\\hat g_t\\mid\\mathcal F_t]=\\nabla f(w_t), $$ $$ \\begin{aligned} \\mathbb E[\\|\\hat g_t-\\nabla f(w_t)\\|^2\\mid\\mathcal F_t] \u0026amp;\\le\\frac{\\sigma^2}{bm} +\\frac{N-m}{m(N-1)}\\cdot\\frac1N\\sum_i \\|\\nabla f_i(w_t)-\\nabla f(w_t)\\|^2\\\\ \u0026amp;\\le\\frac{\\sigma^2}{bm}+\\frac{N-m}{m(N-1)}\\zeta^2. \\end{aligned} $$ The second term is the finite population sampling variance, distinct from the stochastic gradient noise within clients. We derive this coefficient below. Given $\\mathcal F_t$, denote $a_i=\\nabla f_i(w_t)-\\nabla f(w_t)$ and $I_i=\\mathbf 1_{{i\\in S_t}}$; in this case, $\\sum_i a_i=0$. Uniform sampling without replacement satisfies\n$$ \\mathbb E[I_i\\mid\\mathcal F_t]=\\frac mN,\\qquad \\mathbb E[I_iI_j\\mid\\mathcal F_t]=\\frac{m(m-1)}{N(N-1)},\\quad i\\ne j. $$ $$ \\sum_{i\\ne j}\\langle a_i,a_j\\rangle=-\\sum_i\\|a_i\\|^2. $$ Expanding the square and taking expectations separately for diagonal and off-diagonal terms:\n$$ \\begin{aligned} \\mathbb E\\left[\\left\\|\\frac1m\\sum_i I_i a_i\\right\\|^2\\middle|\\mathcal F_t\\right] \u0026amp;=\\frac1{m^2}\\left(\\frac mN-\\frac{m(m-1)}{N(N-1)}\\right)\\sum_i\\|a_i\\|^2\\\\ \u0026amp;=\\frac{N-m}{m(N-1)}\\cdot\\frac1N\\sum_i\\|a_i\\|^2. \\end{aligned} $$ The sample gradient noise has zero mean given $S_t$, so the cross-terms with the client sampling bias here vanish, yielding the aforementioned total variance bound. When $m=N$, the sampling term is zero; with sampling with replacement, the heterogeneity coefficient becomes $1/m$, and the finite population correction for sampling without replacement cannot be applied.\nThus, for $0\u0026lt;\\eta\\le1/L$, directly applying the one-step descent from Part 3 gives\n$$ \\frac1T\\sum_{t\u0026lt;T}\\mathbb E\\|\\nabla f(w_t)\\|^2 \\le\\frac{2\\Delta_0}{\\eta T} +L\\eta\\left(\\frac{\\sigma^2}{bm} +\\frac{N-m}{m(N-1)}\\zeta^2\\right). $$ When $N=1$, one must take $m=1$ and not compute formulas containing $N-1$. For $E\u0026gt;1$, one cannot simply replace the main theorem\u0026rsquo;s $N$ with $m$; one must simultaneously control local drift and client selection.\nExtending the Bound to Multiple Local Steps We still use uniform sampling without replacement, independent of the new sample noise each round. All selected clients start from $w_t$, execute the same $E$ steps, and the server performs an equal-weight average of these $m$ local models. Define\n$$ \\chi_m=\\frac{N-m}{m(N-1)},\\qquad u_t=\\frac1m\\sum_{i\\in S_t}\\nabla f_i(w_t)-h_t, $$ $$ \\begin{aligned} r_t^{(m)}\u0026amp;=\\frac1{mE}\\sum_{i\\in S_t}\\sum_{k=0}^{E-1} \\bigl[\\nabla f_i(w_{i,t}^{(k)})-\\nabla f_i(w_t)\\bigr],\\\\ e_t^{(m)}\u0026amp;=\\frac1{mE}\\sum_{i\\in S_t}\\sum_{k=0}^{E-1}\\varepsilon_{i,t}^{(k)},\\\\ v_t^{(m)}\u0026amp;=h_t+u_t+r_t^{(m)}+e_t^{(m)}. \\end{aligned} $$ $u_t$ is the client sampling error on the common round-initial parameters and cannot be absorbed into the local sample noise. Treating $z_t=u_t+e_t^{(m)}$ as a composite noise for one round, and using the sampling variance and martingale difference properties from the previous subsection, we have\n$$ \\mathbb E[z_t\\mid\\mathcal F_t]=0,\\qquad \\mathbb E[\\|z_t\\|^2\\mid\\mathcal F_t] \\le\\chi_m\\zeta^2+\\frac{\\sigma^2}{bmE}. $$ Here we used $\\mathbb E[e_t^{(m)}\\mid\\mathcal F_t,S_t]=0$, so the cross-terms between $u_t$ and $e_t^{(m)}$ are zero. The cross-terms between drift $r_t^{(m)}$ and $z_t$ are still not assumed to be zero.\nWe also need to verify whether the drift bound is preserved. Define the expected average displacement of the selected clients\n$$ D_{t,k}^{(m)}=\\mathbb E\\left[\\frac1m\\sum_{i\\in S_t} \\|w_{i,t}^{(k)}-w_t\\|^2\\middle|\\mathcal F_t\\right]. $$ At the start of each round, local gradients do not depend on $S_t$, and the selection probability for each client is $m/N$, so\n$$ \\mathbb E\\left[\\frac1m\\sum_{i\\in S_t}\\|\\nabla f_i(w_t)\\|^2 \\middle|\\mathcal F_t\\right] =\\frac1N\\sum_i\\|\\nabla f_i(w_t)\\|^2 \\le\\|h_t\\|^2+\\zeta^2. $$ Therefore, by first expanding the path for selected clients and then taking expectations over selection and noise, $D_{t,k}^{(m)}$ satisfies the same recurrence as 0x03. By Jensen\u0026rsquo;s inequality, we obtain\n$$ \\mathbb E[\\|r_t^{(m)}\\|^2\\mid\\mathcal F_t] \\le8L^2\\eta^2(E-1)^2(\\|h_t\\|^2+\\zeta^2) +\\frac{4L^2\\eta^2(E-1)\\sigma^2}{b}. $$ Now we can repeat the descent proof from 0x04: substitute $z_t$ for $e_t$, preserving its correlation with drift, and replace the noise variance with $\\chi_m\\zeta^2+\\sigma^2/(bmE)$. Define\n$$ A_m=\\frac{L\\eta\\sigma^2}{bm} +L\\eta E\\chi_m\\zeta^2 +\\frac{5L^2\\eta^2(E-1)\\sigma^2}{b} +10L^2\\eta^2(E-1)^2\\zeta^2. $$ defination\nUnder the above uniform without-replacement partial participation model and $0\u003c\\eta\\le1/(8LE)$, $$ \\frac1T\\sum_{t\u0026lt;T}\\mathbb E\\|\\nabla f(w_t)\\|^2 \\le\\frac{4\\Delta_0}{\\eta ET}+4A_m. $$ If the strong convexity condition also holds, then the same $\\rho=1-\\eta E\\mu/2$ yields $$ \\mathcal E_T\\le\\rho^T\\mathcal E_0+\\frac{2A_m}{\\mu}(1-\\rho^T). $$ When $m=N$, $\\chi_m=0$ and $A_m=A$ recover the full participation results. When $E=1$, the structure of \u0026ldquo;noise plus sampling error\u0026rdquo; is also recovered, but the step-size constraints and constants here are more conservative; the precise drift-free analysis from the previous subsection should be prioritized. When $N=m=1$, $\\chi_m=0$ is defined separately.\nThe additional $L\\eta E\\chi_m\\zeta^2$ is a client sampling term, which has a different scale from the original $\\eta^2(E-1)^2\\zeta^2$ local drift term. Increasing $m$ can reduce sampling error, but one cannot thereby claim that local drift also vanishes according to $1/m$. Non-uniform selection depending on training state, consecutive participation of the same batch of clients over multiple rounds, weighted aggregation, and varying numbers of local steps are not directly covered by this extension.\nWeighting Can Change the Target If the goal is $f=\\sum_i p_i f_i$, but in each round a single client is chosen uniformly and its update is fully adopted, then the expected gradient at $E=1$ is $N^{-1}\\sum_i\\nabla f_i$, and not necessarily $\\nabla f$.\nFor example, $p_1=0.9,p_2=0.1$, $f_1(w)=w^2/2$, $f_2(w)=(w-2)^2/2$. The optimum of the weighted objective is $0.2$, while the optimum of the equal-weight objective is $1$. The issue here is not \u0026ldquo;slow convergence,\u0026rdquo; but rather optimizing a different objective.\nIf client $i$ has a selection probability of $\\pi_i\u0026gt;0$, then when computing gradients on the same $w$, the unbiased Horvitz–Thompson form is\n$$ \\hat g(w)=\\sum_{i\\in S}\\frac{p_i}{\\pi_i}g_i(w), \\qquad\\mathbb E[\\hat g(w)]=\\sum_i p_i\\nabla f_i(w). $$ This requires the selection mechanism and new sample noise to satisfy corresponding conditions. Ratio estimators obtained by randomly dividing by \u0026ldquo;the sum of weights of selected clients\u0026rdquo; are generally no longer strictly unbiased; importance weighting may also increase variance. The actual FedAvg aggregation scheme requires specific analysis and cannot be assumed equivalent to the above estimator across all implementations.\nSelecting clients based on online rate, completion speed, or loss may similarly alter the objective and noise conditions. 2 remains applicable here regarding the warning about \u0026ldquo;selecting the fastest clients.\u0026rdquo;\n0x08 Communication, Drift, and Corrections Local Work Is Not Free Progress When the total number of local updates $H=ET$ is fixed, the optimization term in the main bound is $4\\Delta_0/(\\eta H)$. Increasing $E$ can reduce the number of communication rounds, but the drift term increases, and the step size must also satisfy $\\eta LE\\le1/8$. Therefore, when comparing different $E$, one cannot compare only the number of communication rounds or retain only the first term in the bound.\nIf the effective step size $\\alpha=\\eta E$ is fixed, then substituting into $\\eta=\\alpha/E$, the heterogeneity term\u0026rsquo;s $\\eta^2(E-1)^2$ approaches $\\alpha^2$ and will not automatically vanish even if the number of local steps increases indefinitely. For wall-clock time, one must also account for the durations of downloading, uploading, local training, and waiting; this can be combined with the ideas in Efficiency Analysis5, but the homogeneous worker time model in that paper cannot directly represent mobile devices.\nChoosing Parameters from the Bound Step size, number of local steps, and number of participants are not three parameters that can be independently increased without limit. Treating an experiment as a run with fixed $T,E,m,b$, the non-convex bound for partial participation has the following structure:\n$$ B(\\eta)=\\frac{c_0}{\\eta}+c_1\\eta+c_2\\eta^2, \\qquad 0\u0026lt;\\eta\\le\\frac1{8LE}, $$ $$ \\begin{aligned} c_0\u0026amp;=\\frac{4\\Delta_0}{ET},\\\\ c_1\u0026amp;=\\frac{4L\\sigma^2}{bm}+4LE\\chi_m\\zeta^2,\\\\ c_2\u0026amp;=\\frac{20L^2(E-1)\\sigma^2}{b} +40L^2(E-1)^2\\zeta^2. \\end{aligned} $$ If $c_0\u0026gt;0$ and $c_1+c_2\u0026gt;0$, the unconstrained minimum is determined by the unique positive root of the following equation:\n$$ -c_0+c_1\\eta^2+2c_2\\eta^3=0. $$ The left-hand side is strictly increasing on the positive half-axis, so the root can be found via bisection, and then the smaller value with $1/(8LE)$ is taken. When $c_2=0,c_1\u0026gt;0$, we obtain $\\eta=\\sqrt{c_0/c_1}$; when $c_1=c_2=0$, the bound decreases as the step size increases, so the upper limit within the allowed interval suffices. When $c_0=0$, the notion of a unique positive root cannot be applied directly; the residual starting from the optimum must be checked separately.\nThis explains why tuning parameters solely based on $\\eta\\propto1/\\sqrt T$ is not always sufficient: when heterogeneity is high and the number of local steps is large, the quadratic term cannot be ignored. The above formula is used to understand the upper bound, not as a directly copy-pasteable practical optimal learning rate; $L,\\Delta_0,\\sigma^2,\\zeta^2$ is typically unknown, and theoretically allowed step sizes are often conservative.\nFor practical training, one can treat a single local update step as a baseline, then add $E$ and inspect the training curves. When comparing, fix a single budget, such as the total number of local updates or total wall-clock time, and report the number of communication rounds simultaneously; if a set of experiments increases both local computation and communication rounds, comparing only the final accuracy cannot demonstrate whether local updates are more effective.\nDistinguishing Error Sources in Experiments Observation or Modification Primary Quantity Affected in This Model Conclusion That Cannot Be Directly Drawn Increase batch size per step Client-side noise $\\sigma^2/b$ Client distributions become more consistent Increase number of participants per round $m$ Aggregation noise and sampling coefficient $\\chi_m$ Local model drift is eliminated Increase number of local steps $E$ Communication frequency, effective step size, and drift term Convergence speed necessarily improves Decrease local learning rate $\\eta$ Noise residual and drift term, while simultaneously slowing the descent of the optimization term It is necessarily more accurate under a fixed finite budget Adjust aggregation weights Optimized objective $f$ Merely an implementation detail that does not affect the conclusion Furthermore, inconsistent label distributions, inconsistent sample sizes, and inconsistent optima are not the same thing. $\\zeta^2$ measures gradient differences at the parameter location, not parameters for a specific data partition. When partitioning data using a Dirichlet distribution, the concentration parameter can control the partitioning method, but it cannot unconditionally be written as an analytical expression of $\\zeta^2$.\nOne can estimate gradients for each client on the server parameters $w_t$, then observe the differences between gradients; however, differences computed from a finite batch still contain sample noise and cannot be directly treated as true heterogeneity. Repeated sampling, larger diagnostic batches, or explicit noise estimation all help distinguish between the two. In real cross-device training, such additional diagnostics incur communication and data access costs.\nFedProx and SCAFFOLD FedProx6 adds a proximal term to the local objective within a round:\n$$ f_i(w)+\\frac\\lambda2\\|w-w_t\\|^2,\\qquad\\lambda\\ge0. $$ It imposes a cost for local parameters deviating from server parameters. The proximal term cannot guarantee exact elimination of drift in all problems solely based on its form; its analysis also involves objective conditions and local solution accuracy. It is also not the unmodified FedAvg update rule presented in this paper.\n7 uses control variates to correct the direction:\n$$ w_i\\leftarrow w_i-\\eta\\bigl(g_i(w_i)-c_i+c\\bigr). $$ If $c_i=\\nabla f_i(w_t)$ and $c=\\nabla f(w_t)$ are ideally chosen at the start of the round, the mean of the corrected gradients at $w_i=w_t$ exactly equals the global gradient; after local parameters move, there remains error arising from the positional change. The actual algorithm maintains estimates of control variates and must also analyze estimation error, update mechanisms, and sampling. Here we explain the correction mechanism without treating it as a proven SCAFFOLD convergence theorem.\n0x09 What the Analysis Guarantees The main conclusions of this paper are built upon full participation, equal weighting, synchronization, identical local step counts, conditionally unbiased noise, and a unified heterogeneity bound. For non-convex settings, the conclusion is a guarantee of average stationary points; for strongly convex settings, it is geometric contraction toward an error upper bound under a fixed step size. Extensions for partial participation cover uniform without-replacement sampling per round and equal-weight models with identical local step counts; the discussion on weighted sampling addresses objective consistency and does not claim to be a comprehensive theorem covering all FedAvg implementations.\nWhen training different neural networks, one must also check whether smoothness, sampling, batch correlation, participation mechanisms, and optimizers align with the model. Adam, momentum, varying local step counts, asynchronous servers, and non-smooth activations do not automatically satisfy the assumptions of this paper just because the algorithm is also called FedAvg. Verifying accuracy, generalization, fairness, and privacy is also not directly guaranteed by the optimization bounds presented here.\nFor a practical experiment, I would first record the target weights and $N,m,E,b,\\eta$, then separately observe client sampling, local drift, and communication latency. Distinguishing these quantities clearly reveals whether one is improving the optimization direction, reducing stochastic noise, or trading more local computation for less communication.\nReference Sources: 8.\nPart 3\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPart 4\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMcMahan, B. et al. Communication-Efficient Learning of Deep Networks from Decentralized Data, AISTATS 2017. The original algorithm and experiments for FedAvg.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLi, X., Huang, K., Yang, W., Wang, S. and Zhang, Z. On the Convergence of FedAvg on Non-IID Data, ICLR 2020. Strongly convex analysis and fixed step-size issues for non-IID FedAvg.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nEfficiency Analysis\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nLi, T. et al. Federated Optimization in Heterogeneous Networks, MLSys 2020. FedProx and statistical, system heterogeneity.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nKarimireddy, S. P. et al. SCAFFOLD: Stochastic Controlled Averaging for Federated Learning, ICML 2020, especially Sections 2–4. Gradient heterogeneity, local drift, and control variate methods; the more general conditions and tighter bounds in that paper do not equate to the simplified bounds in this paper.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nWoodworth, B., Patel, K. K. and Srebro, N. Minibatch vs Local SGD for Heterogeneous Distributed Learning, NeurIPS 2020. Comparison of local SGD and mini-batch SGD under heterogeneous objectives.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-convergence-analysis-in-deep-learning-part-5/","summary":"\u003ch2 id=\"0x00-preface\"\u003e0x00 Preface\u003c/h2\u003e\n\u003cp\u003ePart 3\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e analyzed mini-batch SGD, Part 4\u003csup id=\"fnref:2\"\u003e\u003ca href=\"#fn:2\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e2\u003c/a\u003e\u003c/sup\u003e discussed synchronous and asynchronous distributed updates. This article turns to federated learning: clients not only hold different data but also perform multiple consecutive updates between a single communication round. Therefore, \u003cstrong\u003elocal stochastic gradients being unbiased for the local objective does not imply that the aggregated direction is unbiased for the current global gradient\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eStarting from FedAvg\u0026rsquo;s update rule, this article provides a simplified proof that can be verified step-by-step: first analyzing the model under \u003cstrong\u003efull participation, equal-weight clients, and identical local steps\u003c/strong\u003e, then deriving extensions for uniform partial participation and discussing inconsistent aggregation weights. The non-convex part guarantees the average gradient norm decreases; the strongly convex part additionally employs the PL inequality. These two conclusions must not be mixed.\u003c/p\u003e","title":"Convergence Analysis in Federated Learning (Part 5)"},{"content":"0x00 Preface This article extends the mini-batch SGD analysis from Part 31 to the distributed setting. We assume $f$ is a $L$-smooth, $\\mu$-strongly convex function on Euclidean space, with an optimal solution existing, $f^{\\ast}=f(w^{\\ast})$. $m$ denotes the number of workers, $b$ the batch size per worker, and $\\sigma^2$ the upper bound on the variance of single-sample gradient noise.\nThe distributed bounds in this paper rely on conditional unbiasedness, noise variance bounds, and corresponding independence assumptions. For scenarios where different workers hold data from different distributions, workers are selected based on completion speed, or computation time depends on the sample, these conditions must be re-verified. Having more machines or faster updates does not, by itself, guarantee that the convergence bounds hold.\n0x01 Distributed Synchronous SGD All workers use the same $w_t$, computing individually\n$$ g_{i,t}=\\frac1b\\sum_{r=1}^b\\nabla\\ell(w_t;\\xi_{i,t,r}),\\qquad \\bar g_t=\\frac1m\\sum_{i=1}^mg_{i,t},\\qquad w_{t+1}=w_t-\\eta\\bar g_t. $$ Given the pre-sampling history $\\mathcal F_t$, we assume the workers\u0026rsquo; gradients are conditionally unbiased with respect to the global objective, their noises are mutually independent, and the noise variance for each worker does not exceed $\\sigma^2/b$. Therefore\n$$ \\mathbb E[\\bar g_t\\mid\\mathcal F_t]=\\nabla f(w_t),\\qquad \\mathbb E[\\|\\bar g_t-\\nabla f(w_t)\\|^2\\mid\\mathcal F_t]\\le\\frac{\\sigma^2}{bm}. $$ defination\nFor $0\u003c\\eta\\le1/L$, let $\\mathcal E_t=\\mathbb E[f(w_t)]-f^{\\ast}$, yielding $$ \\mathcal E_T\\le C_m+(1-\\eta\\mu)^T(\\mathcal E_0-C_m), \\qquad C_m=\\frac{\\eta L\\sigma^2}{2\\mu bm}. $$ The proof simply substitutes $\\sigma^2/(bm)$ into the one-step recurrence from Part 3 and expands the geometric series. With a fixed step size, this guarantees contraction to an error upper bound plateau, not linear convergence to zero error.\nK-Sync and K-Batch-Sync K-Sync accepts gradients from the fastest $K\\le m$ workers and cancels the remaining tasks; K-batch-sync accepts $K$ batches computed on the same parameter, allowing fast workers to contribute multiple batches consecutively.\nIf the batches after selection remain conditionally unbiased, and their noises are conditionally independent or uncorrelated, replace the variance with $\\sigma^2/(bK)$; the error upper bound plateau becomes $C_K=\\eta L\\sigma^2/(2\\mu bK)$. Independence from sampling time and each worker sampling from the same global distribution constitute an ideal model supporting these conditions.\nIf speed correlates with sample categories, selecting the fastest K may introduce bias. Local gradients from non-IID workers are not necessarily unbiased estimates of the global gradient. Thus, one cannot simply substitute $m$ with $K$ while skipping assumption checks.\n0x02 Distributed Asynchronous SGD Asynchronous updates accept a single worker\u0026rsquo;s gradient without waiting for all workers:\n$$ w_{j+1}=w_j-\\eta v_j,\\qquad v_j=g(w_{\\tau_j};\\xi_j),\\qquad0\\le\\tau_j\\le j. $$ $\\tau_j$ is the parameter version read when computing this gradient, and $d_j=j-\\tau_j$ is the staleness step count. $\\xi_j$ represents a batch of size $b$; both the model version and the sampled batch may be random.\nAssumptions for the Proof To ensure the subsequent proof holds step-by-step, we explicitly adopt the following conditions.\nConditional Noise Model. Let $\\mathcal H_j$ contain the server parameter history prior to this update, the selected worker, old parameters, and delay information, but exclude the sample noise of the gradient to be applied. Write $a_j=\\nabla f(w_{\\tau_j})$, $v_j=a_j+e_j$, requiring $$ \\mathbb E[e_j\\mid\\mathcal H_j]=0,\\qquad \\mathbb E[\\|e_j\\|^2\\mid\\mathcal H_j]\\le\\frac{\\sigma^2}{b}. $$ $w_j$ and $w_{\\tau_j}$ are measurable with respect to $\\mathcal H_j$. This condition is more explicit than unbiasedness given only old parameters; if delay or the selection process leaks sample information, a separate proof is required to establish it.\nRelative Gradient Staleness Bound. There exists $0\\le\\gamma\\le1$ such that in every round, $$ \\mathbb E\\|\\nabla f(w_j)-a_j\\|^2 \\le\\gamma\\mathbb E\\|\\nabla f(w_j)\\|^2. $$ This is not a bound on the delay steps $d_j\\le d_{\\max}$, nor does it follow automatically from smoothness. Especially when the current gradient is small, this condition can be quite strong.\nFresh Gradient Probability Lower Bound. Let $\\mathcal P_j$ be the history prior to revealing the delay of the current selection, where $w_j$ is measurable. Assume there exists $p_0\\in[0,1]$ such that $$ \\Pr(\\tau_j=j\\mid\\mathcal P_j)\\ge p_0. $$ Thus\n$$ \\mathbb E\\|a_j\\|^2 \\ge\\mathbb E\\left[\\mathbf1_{\\{\\tau_j=j\\}}\\|\\nabla f(w_j)\\|^2\\right] \\ge p_0\\mathbb E\\|\\nabla f(w_j)\\|^2. $$ This step first applies the probability lower bound given the history, then takes the expectation, avoiding moving the random conditional probability directly outside the expectation.\nConvergence Bound defination\nLet $\\gamma'=1-\\gamma+p_0/2\u003e0$. Under the above conditions and $0\u003c\\eta\\le1/(2L)$, $$ \\mathcal E_T\\le C_{\\rm async}+(1-\\eta\\mu\\gamma\u0026#39;)^T(\\mathcal E_0-C_{\\rm async}), \\qquad C_{\\rm async}=\\frac{\\eta L\\sigma^2}{2\\mu b\\gamma\u0026#39;}. $$ Proof: Take the conditional expectation under $\\mathcal H_j$; the noise cross-terms vanish. Then take the total expectation; smoothness yields\n$$ \\mathbb E[f(w_{j+1})]-\\mathbb E[f(w_j)] \\le-\\eta\\mathbb E\\langle\\nabla f(w_j),a_j\\rangle +\\frac{L\\eta^2}{2}\\mathbb E\\|a_j\\|^2 +\\frac{L\\eta^2\\sigma^2}{2b}. $$ Using $2\\langle u,a\\rangle=\\Vert u\\Vert^2+\\Vert a\\Vert^2-\\Vert u-a\\Vert^2$, we obtain\n$$ \\begin{aligned} \\mathbb E[f(w_{j+1})]-\\mathbb E[f(w_j)] \u0026amp;\\le-\\frac\\eta2(1-\\gamma)\\mathbb E\\|\\nabla f(w_j)\\|^2 -\\frac\\eta2(1-L\\eta)\\mathbb E\\|a_j\\|^2 +\\frac{L\\eta^2\\sigma^2}{2b}\\\\ \u0026amp;\\le-\\frac\\eta2\\left(1-\\gamma+\\frac{p_0}{2}\\right) \\mathbb E\\|\\nabla f(w_j)\\|^2+\\frac{L\\eta^2\\sigma^2}{2b}. \\end{aligned} $$ The second step uses $1-L\\eta\\ge1/2$ and the freshness probability lower bound. Then apply the PL inequality:\n$$ \\mathcal E_{j+1}\\le(1-\\eta\\mu\\gamma\u0026#39;)\\mathcal E_j+\\frac{L\\eta^2\\sigma^2}{2b}. $$ Since $\\mu\\le L$, $\\gamma\u0026rsquo;\\le3/2$, the step size restriction ensures the contraction coefficient lies within $[0,1)$. Expanding the recurrence yields the theorem. When $\\gamma\u0026rsquo;=0$, we cannot divide by it, nor can we derive this geometric contraction bound from it.\nInterpreting Staleness In the idealized model with homogeneous workers, independent non-shifted exponential service times, and instantaneous server processing and model reads, the next completion occurs with equal probability among $m$ workers. For single-gradient asynchronous updates, the freshness probability of the gradient can be taken as $p_0=1/m$.\nAfter stabilization, the lag steps $d_j$ exhibit a geometric tail; the absolute version index $\\tau_j$ does not follow a geometric distribution. The startup phase is constrained by $0\\le d_j\\le j$ and must be considered separately. For heterogeneous times, shifted exponentials, or general service distributions, this probability conclusion cannot be applied directly.\n0x03 K-Async SGD K-Async aggregates gradients from K distinct workers per update; workers completing in this round wait for this update before reading new parameters, while workers not yet completed retain their old computations. The update is\n$$ w_{j+1}=w_j-\\frac\\eta K\\sum_{i=1}^K v_{i,j},\\qquad v_{i,j}=\\frac1b\\sum_{r=1}^b\\nabla\\ell(w_{\\tau_{i,j}};\\xi_{i,j,r}). $$ Given $\\mathcal H_j$ containing all selected old parameters and delays, write $v_{i,j}=a_{i,j}+e_{i,j}$, requiring that the noise be conditionally unbiased, conditionally uncorrelated, and that each variance not exceed $\\sigma^2/b$. Thus, the average noise variance does not exceed $\\sigma^2/(bK)$.\nAdditionally, we require a lower bound on the average relative lag and the freshness probability for each selected gradient:\n$$ \\frac1K\\sum_{i=1}^K\\mathbb E\\|\\nabla f(w_j)-a_{i,j}\\|^2 \\le\\gamma\\mathbb E\\|\\nabla f(w_j)\\|^2, $$ $$ \\Pr(\\tau_{i,j}=j\\mid\\mathcal P_j)\\ge p_0,\\qquad i=1,\\ldots,K. $$ Denote $\\bar a_j=K^{-1}\\sum_i a_{i,j}$. By Jensen\u0026rsquo;s inequality, $\\Vert\\bar a_j\\Vert^2\\le K^{-1}\\sum_i\\Vert a_{i,j}\\Vert^2$. After applying the inner product identity to each $a_{i,j}$ and taking the average, we obtain\n$$ \\begin{aligned} \\mathbb E[f(w_{j+1})]-\\mathbb E[f(w_j)] \u0026amp;\\le-\\frac\\eta2(1-\\gamma)\\mathbb E\\|\\nabla f(w_j)\\|^2 -\\frac\\eta{2K}(1-L\\eta)\\sum_i\\mathbb E\\|a_{i,j}\\|^2 +\\frac{L\\eta^2\\sigma^2}{2bK}\\\\ \u0026amp;\\le-\\frac\\eta2\\gamma\u0026#39;\\mathbb E\\|\\nabla f(w_j)\\|^2 +\\frac{L\\eta^2\\sigma^2}{2bK}. \\end{aligned} $$ Therefore, under the same step-size constraints and $\\gamma\u0026rsquo;\u0026gt;0$ conditions,\n$$ \\mathcal E_T\\le C_{K\\text{-async}}+(1-\\eta\\mu\\gamma\u0026#39;)^T (\\mathcal E_0-C_{K\\text{-async}}),\\qquad C_{K\\text{-async}}=\\frac{\\eta L\\sigma^2}{2\\mu bK\\gamma\u0026#39;}. $$ Here, one must verify $\\gamma$ and $p_0$ for each K separately; they cannot be assumed independent of K, nor can the single-gradient $p_0=1/m$ be blindly copied without conditions. K-batch-async allows a single worker to contribute multiple batches consecutively; it operates under a different runtime model, and the noise and lag conditions must be re-examined.\n0x04 What the Analysis Does Not Guarantee Variance reduction in the synchronous model\u0026rsquo;s $1/m$ and the asynchronous model\u0026rsquo;s $1/K$ both rely on the noise structure. The relative gradient lag bound in the asynchronous setting is an additional condition; merely providing the number of delay steps or a faster update rate cannot substitute for it. These strongly convex conclusions do not directly cover general non-convex neural networks or non-IID federated learning.\nFor combining iteration bounds with stochastic execution times, see Efficiency Analysis in Distributed Machine Learning2. When comparing, one must simultaneously consider batch size, the error upper bound plateau, step size, and the time per round.\nReference This article uses a simplified model with uniformly bounded single-sample variance and explicitly states the conditional noise assumptions required for the proof; the original paper allows more general gradient-dependent variance models, so one cannot omit its additional step-size conditions and directly copy the results.\nSources: 3 4 5.\nPart 3\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nEfficiency Analysis in Distributed Machine Learning\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDutta, S., Joshi, G., Ghosh, S., Dube, P. and Nagpurkar, P., Slow and Stale Gradients Can Win the Race: Error-Runtime Trade-offs in Distributed SGD, AISTATS 2018, especially Theorem 3 and Supplement, Section 8.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDutta, S., Joshi, G., Ghosh, S., Dube, P. and Nagpurkar, P., Slow and Stale Gradients Can Win the Race: Error-Runtime Trade-offs in Distributed SGD, AISTATS 2018, especially Theorem 3 and Supplement, Section 8.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJoshi, G., Optimization Algorithms for Distributed Machine Learning, Springer.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-convergence-analysis-in-deep-learning-part-4/","summary":"\u003ch2 id=\"0x00-preface\"\u003e0x00 Preface\u003c/h2\u003e\n\u003cp\u003eThis article extends the mini-batch SGD analysis from Part 3\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e to the distributed setting. We assume $f$ is a $L$-smooth, $\\mu$-strongly convex function on Euclidean space, with an optimal solution existing, $f^{\\ast}=f(w^{\\ast})$. $m$ denotes the number of workers, $b$ the batch size per worker, and $\\sigma^2$ the upper bound on the variance of single-sample gradient noise.\u003c/p\u003e\n\u003cp\u003eThe distributed bounds in this paper rely on conditional unbiasedness, noise variance bounds, and corresponding independence assumptions. For scenarios where different workers hold data from different distributions, workers are selected based on completion speed, or computation time depends on the sample, these conditions must be re-verified. Having more machines or faster updates does not, by itself, guarantee that the convergence bounds hold.\u003c/p\u003e","title":"Convergence Analysis in Distributed Machine Learning (Part 4)"},{"content":"Preface When discussing privacy-preserving computation, we often say \u0026ldquo;raw data was not uploaded,\u0026rdquo; \u0026ldquo;sensitive attributes have been removed from the representation,\u0026rdquo; or \u0026ldquo;the attack model failed to guess.\u0026rdquo; However, these statements do not directly answer one question: exactly how much more does the attacker know after seeing the system\u0026rsquo;s output?\nInformation theory provides a way to express this question mathematically. It focuses not only on whether data is displayed exactly as-is, but also on whether the attacker\u0026rsquo;s uncertainty has decreased, and under what observations and background knowledge this reduction occurs.\nThis article primarily discusses discrete random variables. For continuous data, mutual information can still be defined via the KL divergence between probability distributions, but differential entropy can be negative, and deterministic mappings may lead to infinite mutual information. Therefore, one cannot mechanically replace every \u0026ldquo;entropy\u0026rdquo; in discrete formulas with its continuous counterpart and assume the result is a directly computable measure of leakage.\nClarifying First: Who Observes What? We describe a basic scenario using four random variables:\nSymbol Meaning A Machine Learning Example $S$ The secret to be protected Sensitive attributes, training set membership, a specific user\u0026rsquo;s data $X$ Data held or processed by the system Images, text, client-side local datasets $Z$ Output visible to the attacker Model predictions, parameters, gradients, or complete communication logs $A$ Auxiliary information the attacker already possesses Candidate samples, public databases, model architecture, other published results A mechanism may transform $X$ into $Z$, but the privacy goal is usually to protect $S$. These two need not be equal: an image may contain content needed for classification as well as identity, location, or other information that should not be leaked.\nDifferent $Z$ correspond to different attack surfaces. Publishing only the final class label differs from publishing the full probability vector; publishing only the final model differs from publishing client gradients for each round. Ignoring $A$ may lead to misjudging a actually recoverable secret as secure.\nEntropy: How Uncertain Were We Before Observation? For a discrete secret $S$, Shannon entropy is defined as\n$$ H(S)=-\\sum_s P(s)\\log_2P(s), $$ By convention, $0\\log_2 0=0$, measured in bits. The entropy of a uniformly random bit is 1; the entropy of a variable that always takes the same value is 0.\nEntropy describes uncertainty in an average sense over the distribution, not the byte count of a data file, nor the sensitivity level. A sensitive attribute that is almost always true may have very low entropy, yet still warrant protection.\nWhen the attacker already knows $A$, a more appropriate baseline is conditional entropy:\n$$ H(S\\mid A)=\\mathbb E_A[H(S\\mid A=a)]. $$ If the auxiliary information already fully determines the secret, then $H(S\\mid A)=0$. In this case, a new publication may not add information, yet this does not mean the secret remains unknown to the attacker.\nFor basic definitions, refer to Stanford\u0026rsquo;s Statistics and Information Theory course notes1.\nMutual Information: How Much More Do We Know After Seeing the Output? Starting from the Reduction in Uncertainty After accounting for auxiliary information, the incremental leakage can be written as\n$$ I(S;Z\\mid A)=H(S\\mid A)-H(S\\mid Z,A). $$ Here, $H(S\\mid Z,A)$ represents the remaining uncertainty after seeing the output. For discrete variables, conditional mutual information is non-negative; under the corresponding probabilistic model, it equals 0 if and only if $S$ and $Z$ are conditionally independent given $A$.\n\u0026ldquo;Mutual information equals 0\u0026rdquo; means the output does not further change the attacker\u0026rsquo;s conditional distribution over the secret; it does not mean the attacker knows nothing, nor does it guarantee security under a different prior or with additional auxiliary information.\nStarting from Changes in Posterior and Prior The same quantity can also be expressed as\n$$ I(S;Z\\mid A)= \\mathbb E_{A,Z}\\left[ D_{\\mathrm{KL}}\\bigl(P_{S\\mid Z,A}\\,\\|\\,P_{S\\mid A}\\bigr) \\right], $$ $$ D_{\\mathrm{KL}}(P\\|Q)=\\sum_sP(s)\\log_2\\frac{P(s)}{Q(s)}. $$ Thus, we can interpret it as: on average, how much does the observation shift the attacker\u0026rsquo;s posterior relative to the prior after having already incorporated the auxiliary information? KL divergence is not a symmetric distance; if $Q(s)=0$ while $P(s)\u0026gt;0$, the corresponding term becomes infinite.\nUnder log loss, the optimal predictor who knows the true conditional distribution has an expected loss equal to the conditional entropy. Therefore, conditional mutual information also corresponds to the average reduction in optimal log loss before and after observation. This interpretation depends on this loss function and cannot be directly converted into a general classification accuracy metric. 2\nAn Example: How Much Does Randomized Response Actually Leak? Let the secret $S$ be a uniformly random bit, the mechanism independently sample $B\\sim\\operatorname{Bernoulli}(q)$, and publish\n$$ Z=S\\oplus B,\\qquad 0\\le q\\le\\frac12. $$ where $\\oplus$ denotes XOR. In other words, the mechanism flips the true answer with probability $q$. Assuming the attacker has no additional relevant information, $Z$ remains a uniformly random bit, and\n$$ I(S;Z)=1-h_2(q), \\qquad h_2(q)=-q\\log_2q-(1-q)\\log_2(1-q). $$ This is because $H(S)=1$, and given $Z$, the uncertainty of the secret is the same as that of the flipped variable. An optimal attacker directly guesses $S=Z$, with a success rate of $1-q$.\nFlip probability $q$ Mutual information, in bits Optimal guessing success rate 0 1 100% 0.1 approx. 0.531 90% 0.25 approx. 0.189 75% 0.5 0 50% Note that even after one quarter of the answer has been flipped, the attacker can still guess correctly with 75% probability. Adding randomness does not mean the secret and the output are independent.\nYou can verify the numbers with the small example below, without needing real personal data:\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 from math import isclose, log2 def binary_entropy(q): if not 0 \u0026lt;= q \u0026lt;= 1: raise ValueError(\u0026#34;q must be between 0 and 1\u0026#34;) if q == 0 or q == 1: return 0.0 return -q * log2(q) - (1 - q) * log2(1 - q) assert isclose(binary_entropy(0), 0.0) assert isclose(binary_entropy(0.5), 1.0) assert isclose(1 - binary_entropy(0.25), 0.18872187554086717) for q in (0, 0.1, 0.25, 0.5): print(q, \u0026#34;leakage bits:\u0026#34;, 1 - binary_entropy(q), \u0026#34;success:\u0026#34;, 1 - q) Why auxiliary information and multiple releases matter? Unconditional mutual information can miss risks Let $R$ be an independent uniform random bit, and let $Z=S\\oplus R$ be released. If the attacker does not know $R$, then $I(S;Z)=0$; but if the auxiliary information is $A=R$, the attacker can directly compute $S=Z\\oplus A$, in which case\n$$ I(S;Z\\mid A)=1. $$ This is not because XOR fails; rather, the condition under which protection holds has changed. When discussing any privacy metric, one must clearly specify what background knowledge the attacker possesses.\nIf individual releases do not leak, joint releases may still leak Continuing with the previous example, release $Z_1=R$ and $Z_2=S\\oplus R$ separately. Viewed individually, both are independent of $S$; viewed together, they allow recovery of $S$.\nIn general, the leakage of the full record $Z_{1:T}$ satisfies the chain rule:\n$$ I(S;Z_{1:T}\\mid A)= \\sum_{t=1}^{T}I(S;Z_t\\mid A,Z_{1:t-1}). $$ Each term on the right must be conditioned on previous release results; one cannot simply sum unconditional single-release mutual informations. This is why federated learning analyses that consider only the gradients of a single round, ignoring prior models and communication content, yield incomplete conclusions.\nWhat does the data processing inequality tell us? If, given $A$, $S\\to X\\to Z$ forms a Markov chain, then\n$$ I(S;Z\\mid A)\\le I(S;X\\mid A). $$ The Markov condition means: once $X$ and $A$ are known, the generation of $Z$ no longer depends additionally on $S$. Therefore, post-processing based solely on existing observations cannot create new information about the secret out of thin air. 3\nBut this is not a proof that \u0026ldquo;compressing makes it secure.\u0026rdquo; If $Z$ retains almost all content related to the secret from $X$, the inequality may still hold while leakage remains large. Conversely, a representation that is easier to use might enable attackers with limited computational power to perform inference more easily, even if its information-theoretic leakage has not increased.\nIf post-processing also accesses new auxiliary information, those data must be included in the model; one cannot continue using a Markov chain that omits the new inputs.\nFano\u0026rsquo;s inequality: from residual uncertainty to guessing error Let the size of the secret\u0026rsquo;s value set be $M\\ge2$, let the attacker construct an estimate $\\widehat S$ based on $(Z,A)$, and let the error probability be $P_e=P(\\widehat S\\ne S)$. Fano\u0026rsquo;s inequality gives\n$$ H(S\\mid Z,A)\\le h_2(P_e)+P_e\\log_2(M-1). $$ If $S$ is uniformly distributed and independent of $A$, and we use $h_2(P_e)\\le1$, we can derive a looser bound:\n$$ P_e\\ge1-\\frac{I(S;Z\\mid A)+1}{\\log_2M}. $$ This shows: for this finite secret space, when leakage is small and the number of candidates is large, the error probability of any estimator has a lower bound. It is not an experimental result of a specific attack algorithm; in binary classification problems, this simplified form usually has no practical constraint and cannot be used to prove that membership inference is close to random guessing. 3\nPrivacy vs. utility: not about removing all information If the output is completely independent of the input, privacy may be excellent, yet the task may become impossible. Let $Y$ be a useful prediction target; one information-theoretic modeling approach is\n$$ \\min_{P_{Z\\mid X}} I(S;Z\\mid A) \\quad\\text{subject to}\\quad I(Y;Z)\\ge u_0. $$ This is a schematic privacy–utility problem: we wish to reduce secret leakage while retaining task-relevant information. Practical systems may define utility using accuracy, reconstruction error, or other metrics, rather than directly using $I(Y;Z)$. If $S$ is strongly correlated with $Y$, these two objectives may inherently conflict. 2\nA common experimental approach is to use a neural-network-based attacker to estimate sensitive attributes, then train an encoder to make the attacker fail. However, the cross-entropy of a finite model is typically larger than the optimal conditional entropy; an underpowered attacker will also yield high loss. Therefore, making an attacker fail does not prove that the true mutual information is small.\nWhat is the difference between information-theoretic privacy and differential privacy? Differential Privacy (DP) compares the output distributions of adjacent datasets. For a specified adjacency relation, and for all adjacent $D,D\u0026rsquo;$ and output events $E$, it is required that\n$$ P(\\mathcal M(D)\\in E) \\le e^{\\varepsilon}P(\\mathcal M(D\u0026#39;)\\in E)+\\delta. $$ The logarithm here uses the natural base, whereas the information quantity mentioned earlier uses $\\log_2$. $\\varepsilon$ constrains the distinguishability of the output distribution, while $\\delta$ is the additive relaxation term for approximate DP and cannot be directly interpreted as \u0026ldquo;the probability that the system leaks all data is $\\delta$\u0026rdquo;. The definition also needs to clarify whether adjacent datasets differ by adding/deleting one record, replacing one record, or adding/deleting all records of a user. 4\nDimension Conditional Mutual Information Perspective DP Perspective Core Question How much secret uncertainty is reduced on average in the output How much the output distribution can change given a change in protected data Objects to Clarify Secret, joint distribution, auxiliary information, and observation Random mechanism, adjacency relation, and privacy parameters Form of Guarantee Average leakage amount under a given probabilistic model Holds for all adjacent inputs and output events under a specified adjacency relation Common Misuse Mistaking an attacker\u0026rsquo;s failure for zero mutual information Adding noise without analyzing sensitivity, sampling, and cumulative budget There is a theoretical connection between the two, but they are not the same definition. A small mutual information under a given prior does not automatically imply DP with specified parameters; nor does DP require that the model not learn any overall statistical patterns.\nThe aforementioned binary randomized response mechanism satisfies pure local DP on single-bit inputs when $0\u0026lt;q\\le1/2$, with parameters\n$$ \\varepsilon=\\ln\\frac{1-q}{q}. $$ This is obtained by directly comparing the two output probabilities corresponding to the two inputs. When $q=1/2$, $\\varepsilon=0$; when $q\\to0$, the privacy parameter tends to infinity. Its DP parameters do not depend on a uniform prior, whereas the specific value of $I(S;Z)=1-h_2(q)$ mentioned earlier depends on the assumption of a uniform secret.\nHow to Use These Concepts in Machine Learning Experiments First specify the secret, then specify the complete content visible to the attacker. For example, membership status, a specific attribute, and an entire training sample are different secrets and cannot be uniformly replaced by a single reconstruction error.\nSubsequently, distinguish three levels: the distribution and mechanism analyzed theoretically, the mutual information or other proxy quantities estimated numerically, and the results obtained by a specific attacker implementation. These three can provide evidence for one another but cannot be arbitrarily interchanged.\nEstimating mutual information for high-dimensional continuous data is particularly difficult. Discretization, sample size, estimator bias, and the training process all affect the results; numerical values obtained under different discretization schemes should not be directly compared horizontally. In federated learning, one must also incorporate multi-round interactions and auxiliary information into the analysis, rather than selecting only a single update that appears \u0026ldquo;non-leaking\u0026rdquo;.\nConclusion Information theory shifts the privacy question from \u0026ldquo;whether raw data was sent\u0026rdquo; to \u0026ldquo;how much more is known after observation\u0026rdquo;. It is particularly well-suited for explaining average leakage, the role of auxiliary information, and the relationship with multiple releases.\nHowever, an information quantity formula itself is not a complete security proof. Clearly specifying the secret, distribution, attack surface, and assumptions, and distinguishing between theoretical bounds, estimated values, and actual attack results, is key to applying these concepts in research.\nReferences and Further Reading Sources: 5.\nStanford：Lecture Notes on Statistics and Information Theory。\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMakhdoumi et al.: From the Information Bottleneck to the Privacy Funnel.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMIT：Information Theory，Lecture 2。\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDwork and Roth: The Algorithmic Foundations of Differential Privacy.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPrivacy Leaks in Deep Learning: Examining specific attack surfaces from membership inference, gradient reconstruction, and training data extraction.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-information-theory-for-privacy/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eWhen discussing privacy-preserving computation, we often say \u0026ldquo;raw data was not uploaded,\u0026rdquo; \u0026ldquo;sensitive attributes have been removed from the representation,\u0026rdquo; or \u0026ldquo;the attack model failed to guess.\u0026rdquo; However, these statements do not directly answer one question: exactly how much more does the attacker know after seeing the system\u0026rsquo;s output?\u003c/p\u003e\n\u003cp\u003eInformation theory provides a way to express this question mathematically. It focuses not only on whether data is displayed exactly as-is, but also on whether the attacker\u0026rsquo;s uncertainty has decreased, and under what observations and background knowledge this reduction occurs.\u003c/p\u003e","title":"Information Theory in Privacy-Preserving Computation: How to Understand and Measure Information Leakage"},{"content":"Preface After model training is complete, is the training data safe? If training occurs on the client side and the server only receives parameter updates, is privacy already protected?\nThe answers to these questions depend on what the attacker can observe. While raw data may not be sent directly, prediction probabilities, model parameters, and gradients can still contain information related to the training data.\nStarting from specific attack objectives, this article introduces several types of privacy leakage in deep learning, explaining the conditions required, how to evaluate them, and which parts of the data common defenses actually protect. The numerical examples below use artificially constructed data and do not involve real personal information.\nEstablishing the Threat Model First At minimum, four things must be clarified: who the attacker is, what they can observe, what auxiliary information they know, and which secret they hope to recover.\nScenario Attacker\u0026rsquo;s Observation Possible Additional Knowledge User of a classification service Final label, confidence, or full probability vector Candidate samples and their true labels, data from the same distribution Person obtaining model files Parameters, structure, intermediate representations, or gradients Training pipeline, preprocessing methods, partial training data Federated learning server Models from each round, client updates, or aggregation results Client identities, participation records, optimizer configurations Collaborating clients Received global model and their own local updates Data they hold, some other publicly available information Black-box and white-box are not the only two scenarios. An API that returns only a label and one that provides a probability vector may both be called black-box, yet they differ significantly in information content.\nOne must also distinguish between honest-but-curious attackers and actively malicious attackers: the former analyze observations after executing the protocol as specified, while the latter may alter the models sent, participant selection, or other protocol steps. Conclusions valid for the former cannot be directly extended to the latter.\nWhat Are the Different Attack Types Asking? Attack Type Objective Objects Not to Confuse Membership Inference Determine whether a candidate sample participated in training Does not equal recovering the full content of that sample Attribute Inference Infer sensitive attributes of a sample, user, or training set Does not necessarily require knowing exact membership status Model Inversion / Gradient Reconstruction Recover input information from model outputs, representations, or gradients Representative images of a class are not necessarily real training samples Training Data Extraction Recover actual training content from the model Distinguish between ordinary generation, similar content, and verbatim reproduction Model stealing primarily focuses on replicating model capabilities or parameters, which is a different objective from personal data privacy; adversarial examples primarily focus on whether predictions are manipulated and cannot be directly used as privacy attack metrics.\nMembership Inference: Why Do Predictions Reveal Training Identity? A Simple Loss Threshold Attack Let the candidate sample be $(x,y)$. The attacker can obtain the target model\u0026rsquo;s prediction and knows its label. One of the simplest approaches is to compute the loss and make a judgment based on a threshold:\n$$ \\widehat m(x,y)=\\mathbf 1\\{\\ell(f_w(x),y)\\le\\tau\\}, $$ where $m=1$ denotes a training set member. This attack relies on an empirical signal: models often fit training samples better, so members may have lower loss.\nHowever, \u0026ldquo;low loss\u0026rdquo; and \u0026ldquo;membership status\u0026rdquo; are not the same thing. Non-members that are easy to classify may also have very low loss, while difficult training samples may have very high loss. The threshold should be selected on independent calibration data; if the test data used for the final report is also used to tune the threshold, the attack capability will be overestimated.\nWork by Shokri et al. uses shadow models and an attack classifier to learn the prediction differences between members and non-members. It demonstrates privacy risks in black-box outputs, but this does not mean every model or every dataset is equally vulnerable to attack. 1\nWhy Is Average Accuracy Insufficient? Definition\n$$ \\operatorname{TPR}=P(\\widehat m=1\\mid m=1), \\qquad \\operatorname{FPR}=P(\\widehat m=1\\mid m=0). $$ TPR is the proportion of true members correctly identified, and FPR is the proportion of non-members incorrectly judged as members. If the proportion of true members among candidate samples is $\\pi$, then the precision of a \u0026ldquo;member\u0026rdquo; judgment is\n$$ P(m=1\\mid\\widehat m=1)= \\frac{\\pi\\operatorname{TPR}} {\\pi\\operatorname{TPR}+(1-\\pi)\\operatorname{FPR}}. $$ For example, when $\\pi=1%$, TPR is 80%, and FPR is 1%, the precision is only about 44.7%. Among 10,000 candidates, approximately 80 true members are identified, yet 99 non-members are also falsely reported.\nTherefore, reporting a single accuracy on a test set where members and non-members each account for half does not fully reflect real-world scenarios. Work by Carlini et al. on LiRA emphasizes examining identification capability under low false positive rates, such as TPR@0.1% FPR. 2\nEstimating a low false positive rate itself requires a sufficient number of non-member samples. Observing zero false positives on a test set with only a few hundred non-members does not justify claiming the actual FPR is 0, nor does it allow for stable evaluation of false positive rates at the one-in-a-thousand level.\nOverfitting Is Not the Only Explanation Overfitting can provide a signal for membership inference, but having training and test average losses close together does not prove that every sample is safe. Averages can also mask a small number of strongly memorized samples, especially those that are repetitive, rare, or anomalous.\nWhen evaluating, try to ensure that members and non-members come from matched data distributions, and control differences in categories, preprocessing, and data collection sources. Otherwise, the attacker may simply be distinguishing between two datasets rather than inferring membership.\nAttribute Inference and Model Inversion Attribute inference does not need to reconstruct the entire sample. For instance, an attacker may already know some features of a person and wish to infer another attribute from the model output; or they may infer the data distribution characteristics of a client from its client updates.\nHere, we must distinguish between two types of information: the general patterns learned by the model, and the information specific to a particular training record. A model inferring a sensitive attribute from public features may already have real-world privacy implications, but one cannot claim it memorized that person\u0026rsquo;s training record solely based on successful inference.\nModel inversion typically attempts to find inputs that explain a model\u0026rsquo;s output or intermediate representation. An image that a classifier identifies with high confidence as belonging to a specific person might merely reflect a representative pattern preferred by the model. To claim that real training data has been recovered, one must verify against training records, report matching criteria and reconstruction quality, rather than simply displaying results that \u0026ldquo;look like\u0026rdquo; the target.\nWhy Gradients Might Leak Inputs? Analytical Example with a Single-Sample Linear Model Let\u0026rsquo;s start with the simplest model:\n$$ f_{w,b}(x)=w^{\\mathsf T}x+b, \\qquad \\ell=\\frac12(f_{w,b}(x)-y)^2. $$ Let the residual be $r=f_{w,b}(x)-y$, then\n$$ g_w=\\nabla_w\\ell=rx, \\qquad g_b=\\frac{\\partial\\ell}{\\partial b}=r. $$ If an attacker observes these two single-sample gradients and $g_b\\ne0$, they can divide coordinate by coordinate:\n$$ x=\\frac{g_w}{g_b}. $$ The original input was not uploaded, yet the gradients are sufficient to recover it. Below, we verify this using only three dimensions of synthetic data:\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 import numpy as np x = np.array([0.2, 0.7, -0.4]) w = np.array([0.3, -0.2, 0.5]) b, y = 0.1, 1.0 residual = w @ x + b - y weight_gradient = residual * x bias_gradient = residual assert bias_gradient != 0 reconstructed = weight_gradient / bias_gradient assert np.allclose(reconstructed, x) print(reconstructed) This is not a formula that holds for all models or all batch settings. If $g_b=0$, division is not applicable; if one observes the average gradient of multiple samples, the numerator and denominator become weighted sums, and typically one cannot recover each input using this ratio.\nFrom Analytical Recovery to Gradient Matching For neural networks, reconstruction can be understood as searching for a set of candidate inputs whose generated gradients are close to the observed gradients:\n$$ \\min_{\\widetilde x,\\widetilde y} \\left\\|\\nabla_w\\ell(f_w(\\widetilde x),\\widetilde y)-g_{\\mathrm{obs}}\\right\\|^2 +\\lambda R(\\widetilde x). $$ where $R$ represents a prior or regularization term; this expression is merely illustrative for the single-sample case. Deep Leakage from Gradients demonstrates the possibility of recovering training inputs via gradient matching; subsequent work, Inverting Gradients, used direction-dependent objectives and stronger optimization strategies to further investigate image reconstruction. 3, Inverting Gradients4\nPractical difficulty is influenced by conditions such as batch size, model architecture, parameter state, label knowledge, preprocessing, normalization information, and whether observations include multi-step updates. Failure to reconstruct under a specific setting only indicates that the attack implementation was unsuccessful; increasing the batch size, compressing, or clipping gradients does not automatically constitute a rigorous privacy guarantee.\nTraining Data Extraction: Do Models Really Recite the Original Text? Membership inference asks \u0026ldquo;was this data used?\u0026rdquo;, while training data extraction asks \u0026ldquo;can the actual content used be retrieved?\u0026rdquo; Language models may generate fragments identical to training content, but one must distinguish between general knowledge, common expressions, content obtainable from public sources, and verifiable reproductions of training samples.\nResearch by Carlini et al. extracted and verified training text from language models, demonstrating the risks of model memorization and training data recovery. 5\nWhen studying such risks, one can use artificially generated, unique, and non-sensitive canaries as controlled objects, recording whether they were added to training, their repetition counts, and the model\u0026rsquo;s generation results. However, canary experiments measure memorization behavior under specific experimental conditions and cannot fully replace risk assessments for real data.\nThe fact that a model does not directly recite content does not mean that membership, attributes, or other information have not leaked. Different attack objectives require separate evaluation.\nWhy Federated Learning Does Not Automatically Solve Privacy Issues? Federated learning changes the location of data processing: raw data remains on the client, and updates are sent after local training. This architecture can reduce centralized collection of raw data, but the updates themselves may still serve as carriers of sensitive information.\nIf the server can see updates from individual clients, it can directly analyze these updates; if it only sees aggregated results from multiple clients, the attack surface changes, but one must still consider the number of participants, collusion, cross-round information, and whether the server can actively alter training conditions.\nSecure Aggregation uses cryptographic protocols to allow the server to obtain the aggregated value under corresponding security assumptions without directly obtaining the plaintext inputs of individual participants. It protects the visibility of individual inputs during the computation process but does not require the aggregated result or the final model to be completely independent of individual data. 6\nTherefore, Secure Aggregation and Differential Privacy (DP) address different problems and can be used in combination. One must also match assumptions regarding the number of participants, dropouts, and collusion thresholds specific to the protocol; one cannot simply write \u0026ldquo;aggregation was used\u0026rdquo; and treat ordinary averaging as cryptographic secure aggregation.\nAdditionally, how client data is partitioned affects the research scenario. When constructing Non-IID data via Dirichlet Distribution7, some clients may have very few samples or highly concentrated labels; however, these statistical phenomena alone do not constitute proof of a specific attack\u0026rsquo;s success, and explicit observation and experimental conditions are still required.\nWhat Do Common Defenses Protect? Method Primary Effect Boundaries to Retain Data minimization, sensitive field handling, and deduplication Reduces sensitive content entering the system before training and opportunities for repeated memorization Cannot guarantee by itself that remaining data will not leak Training measures such as regularization and early stopping Improves generalization and may reduce some membership inference signals Good average generalization does not equate to per-sample privacy guarantees Reduce probability outputs, query limits, and access control Limit attackers\u0026rsquo; observation and invocation capabilities Must cover model files, logs, and other accessible outputs Clipping, quantization, compression, or larger batch sizes Change the information carried by updates and the difficulty of attacks Cannot claim a DP guarantee without privacy analysis Secure aggregation, encrypted computation, or trusted execution environments Protect the computation process under specified trust and protocol assumptions Released outputs may still enable inference Differential Privacy Limit output distribution changes caused by altering protected units Adjacency relations, sampling methods, and total privacy budget must be explicitly defined DP-SGD is not just about adding a bit of noise to gradients The core of DP-SGD includes clipping gradients per sample and adding calibrated noise after aggregation. A common schematic form is\n$$ \\overline g_i=\\frac{g_i}{\\max(1,\\|g_i\\|_2/C)}, \\qquad \\widetilde g=\\frac1{|\\mathcal B|} \\left(\\sum_{i\\in\\mathcal B}\\overline g_i+\\xi\\right), \\qquad \\xi\\sim\\mathcal N(0,\\sigma^2 C^2 I). $$ $C$ is the clipping threshold, $\\sigma$ is the noise multiplier, and $\\mathcal B$ is the current batch. This illustrates the algorithmic structure; specific sensitivity and $(\\varepsilon,\\delta)$ calculations depend on adjacency definitions, sampling mechanisms, and privacy accounting methods, and cannot be directly read from this formula alone. 8\nClipping only the batch gradient does not justify reusing the sensitivity bound or privacy accounting for per-sample clipping; adding noise and then providing unnoised gradients to an attacker does not protect that additional observation via public DP guarantees. Optimizers, checkpoints, data-dependent hyperparameter tuning, or additional statistical releases must also be incorporated into the analysis.\nIn federated learning, one must distinguish between record-level and user-level protection. If a user contributes many records, record-level DP parameters cannot be directly applied as user-level DP parameters. User-level mechanisms typically require controlling influence at a level matching user contributions, combined with analysis of actual participation and release methods.\nDP does not require hiding all group patterns, nor does it guarantee that no sensitive attribute can be inferred from existing public information. For further discussion on adjacency relations, auxiliary information, and information content, refer to Information Theory in Privacy-Preserving Computation9.\nHow to Design a Persuasive Privacy Experiment? First, clearly define the observation scope and the protected unit. An attack experiment providing only final predictions cannot represent the risk of releasing all training gradients; studying a single record does not directly demonstrate the security of all data for a user.\nThen, select corresponding metrics for different goals: report TPR at low FPR and prior assumptions for membership inference; report matching, error, and success rates against real inputs for reconstruction attacks; report verifiable matching criteria and counts for data extraction attacks. Do not let a single impressive reconstruction image replace statistical analysis of the entire dataset.\nCalibration and final evaluation data should be separated, and one must check whether the distributions of members and non-members match. Training, attack optimization, and partitioning random seeds also affect results; fluctuations and failure cases should be reported, not just selected successful images.\nFinally, evaluate defenses using multiple reasonable attack baselines, considering whether the attacker knows the defense settings. Theoretical guarantees require proven or trusted-implementation privacy accounting; empirical risk requires actual attack assessment. Attack failures are useful evidence but cannot replace strict guarantees.\nConclusion Privacy leakage in deep learning is not a single problem. Knowing whether a sample participated in training, inferring a specific sensitive attribute, reconstructing inputs from gradients, and extracting memorized training content involve different secrets and different attack surfaces.\nOnly after understanding these distinctions can we find the correct placement for defenses: which information should never enter training, which intermediate results need hiding, which released results need limiting individual influence, and what experimental results actually prove. Protecting privacy requires not just \u0026rsquo;not transmitting raw data,\u0026rsquo; but also a complete and explicit threat model.\nReferences and Further Reading Shokri et al.: Membership Inference Attacks against Machine Learning Models.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCarlini et al.: Membership Inference Attacks From First Principles.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nZhu, Liu, and Han: Deep Leakage from Gradients.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGeiping et al.: Inverting Gradients.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCarlini et al.: Extracting Training Data from Large Language Models.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBonawitz et al.: Practical Secure Aggregation for Privacy-Preserving Machine Learning.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nDirichlet Distribution\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nAbadi et al.: Deep Learning with Differential Privacy.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nInformation Theory in Privacy-Preserving Computation\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-privacy-leakage-in-deep-learning/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eAfter model training is complete, is the training data safe? If training occurs on the client side and the server only receives parameter updates, is privacy already protected?\u003c/p\u003e\n\u003cp\u003eThe answers to these questions depend on what the attacker can observe. While raw data may not be sent directly, prediction probabilities, model parameters, and gradients can still contain information related to the training data.\u003c/p\u003e\n\u003cp\u003eStarting from specific attack objectives, this article introduces several types of privacy leakage in deep learning, explaining the conditions required, how to evaluate them, and which parts of the data common defenses actually protect. The numerical examples below use artificially constructed data and do not involve real personal information.\u003c/p\u003e","title":"Privacy Leakage in Deep Learning: What Models, Predictions, and Gradients Reveal"},{"content":"0x00 Preface This article analyzes mini-batch SGD, following the Euclidean norm and stochastic gradient definitions from Part 11, and comparing them with the deterministic gradient descent from Part 22.\nLet $\\mathcal F_t$ denote the history before the $t$-th sampling round, using a fixed step size to update $w_{t+1}=w_t-\\eta g_t$. $g_t$ is the average of $b$ new sample gradients, assuming\n$$ \\mathbb E[g_t\\mid\\mathcal F_t]=\\nabla f(w_t),\\qquad \\mathbb E[\\|g_t-\\nabla f(w_t)\\|^2\\mid\\mathcal F_t]\\le\\frac{\\sigma^2}{b}. $$ Here $\\sigma^2$ is the upper bound on the variance of single-sample noise; the reduction by a factor of $1/b$ relies on conditionally independent noise within the batch or zero cross-covariance.\nOne-step Descent $L$-smoothness yields\n$$ f(w_{t+1})\\le f(w_t)-\\eta\\langle\\nabla f(w_t),g_t\\rangle +\\frac{L\\eta^2}{2}\\|g_t\\|^2. $$ First, take the expectation conditioned on $\\mathcal F_t$, where $w_t$ is fixed:\n$$ \\mathbb E[f(w_{t+1})\\mid\\mathcal F_t] \\le f(w_t)-\\eta\\left(1-\\frac{L\\eta}{2}\\right)\\|\\nabla f(w_t)\\|^2 +\\frac{L\\eta^2\\sigma^2}{2b}. $$ Next, take the expectation over the history. If $0\u0026lt;\\eta\\le1/L$, we obtain the subsequent recursive relation shared by all:\n$$ \\mathbb E[f(w_{t+1})]\\le\\mathbb E[f(w_t)] -\\frac\\eta2\\mathbb E\\|\\nabla f(w_t)\\|^2 +\\frac{L\\eta^2\\sigma^2}{2b}. $$ One cannot treat the random $f(w_t)$ or gradient as deterministic quantities in formulas involving unconditional expectations.\n0x01 Smooth and Strongly Convex with Mini-batch SGD defination\nIf $f$ is $L$-smooth, $\\mu$-strongly convex, and an optimal solution exists, define $$ \\mathcal E_t=\\mathbb E[f(w_t)]-f^*,\\quad q=1-\\eta\\mu,\\quad C_b=\\frac{\\eta L\\sigma^2}{2\\mu b}. $$ For $0\u003c\\eta\\le1/L$, we have $$ \\mathcal E_T\\le q^T\\mathcal E_0+C_b(1-q^T) =C_b+q^T(\\mathcal E_0-C_b). $$ Proof: By the PL condition $\\Vert\\nabla f(w_t)\\Vert^2\\ge2\\mu(f(w_t)-f^{\\ast})$, a single step of descent becomes\n$$ \\mathcal E_{t+1}\\le(1-\\eta\\mu)\\mathcal E_t+\\frac{L\\eta^2\\sigma^2}{2b}. $$ Since $0\\le q\u0026lt;1$, expanding the recursion and summing the geometric series yields\n$$ \\mathcal E_T\\le q^T\\mathcal E_0 +\\frac{L\\eta^2\\sigma^2}{2b}\\sum_{s=0}^{T-1}q^s =q^T\\mathcal E_0+C_b(1-q^T). $$ What Does the Bound Guarantee? Under a fixed step size, the transient term in the bound decreases at a geometric rate, but typically only guarantees $\\limsup_T\\mathcal E_T\\le C_b$. This is an upper bound on the error, not the actual error the algorithm necessarily achieves, nor a guarantee of linear convergence to the exact optimal solution.\nIf $0\u0026lt;q\u0026lt;1$ and $\\mathcal E_0\u0026gt;\\epsilon\u0026gt;C_b$, the required number of iterations is\n$$ T\\ge\\left\\lceil \\frac{\\log((\\mathcal E_0-C_b)/(\\epsilon-C_b))}{-\\log q} \\right\\rceil. $$ If $\\epsilon\\le C_b$, this upper bound cannot guarantee reaching the target. One can reduce the step size, increase the batch size, or use a decaying step size scheme with applicable conditions. When $\\sigma^2=0$, the results of deterministic GD are recovered; when $q=0$, directly use the recursion without calculating $\\log0$.\n0x02 Smooth Non-convex with Mini-batch SGD defination\nAssume $f$ is $L$-smooth and has a finite lower bound $f_{\\inf}$, and the stochastic gradients satisfy the above conditions. For $0\u003c\\eta\\le1/L$ and $T\\ge1$, $$ \\frac1T\\sum_{t=0}^{T-1}\\mathbb E\\|\\nabla f(w_t)\\|^2 \\le\\frac{2(f(w_0)-f_{\\inf})}{\\eta T}+\\frac{L\\eta\\sigma^2}{b}, $$ Here we assume the starting point $w_0$ is deterministic; for a random starting point, replace $f(w_0)$ with its expectation. Proof: Summing the one-step descent from $t=0$ to $T-1$ yields\n$$ \\frac\\eta2\\sum_{t=0}^{T-1}\\mathbb E\\|\\nabla f(w_t)\\|^2 \\le f(w_0)-\\mathbb E[f(w_T)]+\\frac{TL\\eta^2\\sigma^2}{2b}. $$ Using $f(w_T)\\ge f_{\\inf}$, then dividing by $\\eta T/2$ gives the conclusion. This summation starts from $w_0$, so the telescoping endpoints are $f(w_0)$ and $f(w_T)$.\nStep Size and Output If $R$ are independent and uniformly distributed over $\\lbrace 0,\\ldots,T-1\\rbrace$, the expected squared gradient of the random output $w_R$ equals the above average. This conclusion cannot be directly converted into a guarantee for the final point $w_T$, nor does it guarantee reaching the global optimum.\nDenote $A=f(w_0)-f_{\\inf}\u0026gt;0$. When $\\sigma^2\u0026gt;0$, choose for a given run length $T$\n$$ \\eta=\\min\\left\\{\\frac1L,\\sqrt{\\frac{2Ab}{L\\sigma^2T}}\\right\\}. $$ When the second term does not exceed $1/L$, the bound on the average squared gradient is\n$$ 2\\sqrt{\\frac{2AL\\sigma^2}{bT}}. $$ This is a fixed step size chosen based on a predetermined run length, not replacing the current index with $T$ at every step. It demonstrates the $O(1/\\sqrt{bT})$ behavior within the corresponding range; the step size upper limit constrains this scaling. Without noise, taking $\\eta=1/L$ yields $2LA/T$.\nReference Sources: 3 4.\nPart 1\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nPart 2\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBottou, L., Curtis, F. E. and Nocedal, J., Optimization Methods for Large-Scale Machine Learning, Section 4, especially Theorems 4.6 and 4.8.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJoshi, G., Optimization Algorithms for Distributed Machine Learning, Springer.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-convergence-analysis-in-deep-learning-part-3/","summary":"\u003ch2 id=\"0x00-preface\"\u003e0x00 Preface\u003c/h2\u003e\n\u003cp\u003eThis article analyzes mini-batch SGD, following the Euclidean norm and stochastic gradient definitions from Part 1\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e, and comparing them with the deterministic gradient descent from Part 2\u003csup id=\"fnref:2\"\u003e\u003ca href=\"#fn:2\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e2\u003c/a\u003e\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eLet $\\mathcal F_t$ denote the history before the $t$-th sampling round, using a fixed step size to update $w_{t+1}=w_t-\\eta g_t$. $g_t$ is the average of $b$ new sample gradients, assuming\u003c/p\u003e\n\u003cdiv class=\"has-mathjax\"\u003e\n\n$$\n\\mathbb E[g_t\\mid\\mathcal F_t]=\\nabla f(w_t),\\qquad\n\\mathbb E[\\|g_t-\\nabla f(w_t)\\|^2\\mid\\mathcal F_t]\\le\\frac{\\sigma^2}{b}.\n$$\n\n\u003c/div\u003e\n\u003cp\u003eHere $\\sigma^2$ is the upper bound on the variance of single-sample noise; the reduction by a factor of $1/b$ relies on conditionally independent noise within the batch or zero cross-covariance.\u003c/p\u003e","title":"Convergence Analysis in Deep Learning (Part 3)"},{"content":"0x00 Preface This post analyzes deterministic first-order optimization methods. Basic definitions are in Part 11. We default to the Euclidean norm and distinguish between gradient descent and the subgradient method for non-smooth problems.\nThe smooth section discusses unconstrained problems; the non-smooth section allows a closed convex feasible set $\\mathcal C$, using projections to keep iterates within the feasible set. Assume an optimal solution $w^{\\ast}$ exists, $f^{\\ast}=f(w^{\\ast})$.\n0x01 Definition of Convergence First, we must specify what error we are measuring:\nFunction value error: $f(w_T)-f^{\\ast}$. Distance to the optimal solution: $\\Vert w_T-w^{\\ast}\\Vert^2$; if there are multiple optimal solutions, one can also discuss the distance to the set of optimal solutions. Stationarity measure in non-convex problems: $\\min_{0\\le t\u0026lt;T}\\Vert\\nabla f(w_t)\\Vert^2$, or the average of squared gradient norms. These metrics cannot be arbitrarily interchanged. A small gradient norm does not imply finding a global optimum, nor does it alone guarantee that the iteration sequence converges to a specific point.\nLinear convergence typically refers to an error upper bound of $Cq^T$, $0\u0026lt;q\u0026lt;1$; $O(1/T)$ and $O(1/\\sqrt T)$ belong to sublinear convergence. For a positive error sequence $e_t$, Q-superlinear convergence means $e_{t+1}/e_t\\to0$, while Q-quadratic convergence requires that eventually $e_{t+1}\\le Ce_t^2$. We do not use $\\log\\log e_t$ to describe quadratic convergence because when $e_t\u0026lt;1$, this real-valued expression is undefined.\nThe complexities below refer to the number of iterations required to guarantee that the corresponding error does not exceed $\\epsilon$; constants depend on the objective function and the starting point.\n0x02 Smooth Strongly Convex with GD defination\nIf $f$ is $L$-smooth and $\\mu$-strongly convex, using a fixed step size $0\u003c\\eta\\le1/L$, then $$ f(w_T)-f^*\\le(1-\\eta\\mu)^T(f(w_0)-f^*). $$ By the descent lemma and the update $w_{t+1}=w_t-\\eta\\nabla f(w_t)$:\n$$ \\begin{aligned} f(w_{t+1})-f(w_t) \u0026amp;\\le-\\eta\\left(1-\\frac{L\\eta}{2}\\right)\\|\\nabla f(w_t)\\|^2\\\\ \u0026amp;\\le-\\frac\\eta2\\|\\nabla f(w_t)\\|^2 \\le-\\eta\\mu(f(w_t)-f^*). \\end{aligned} $$ The final step uses the PL inequality. Rearranging yields a single-step recurrence; iterating $T$ times suffices.\nWhen $\\eta=1/L$, $\\mu\u0026lt;L$, setting $E_0=f(w_0)-f^{\\ast}$ gives $f(w_T)-f^{\\ast}\\le e^{-\\mu T/L}E_0$, so $T\\ge(L/\\mu)\\log(E_0/\\epsilon)$ is sufficient. The condition number is $L/\\mu$; this result exhibits linear convergence. If $\\mu=L$, the above recurrence yields zero error in a single step.\n0x03 Smooth Convex with GD defination\nIf $f$ is convex and $L$-smooth, using $\\eta=1/L$, then for $T\\ge1$, $$ f(w_T)-f^*\\le\\frac{L\\|w_0-w^*\\|^2}{2T}. $$ Let $R_t=\\Vert w_t-w^{\\ast}\\Vert$. By convexity and the distance identity of the update,\n$$ f(w_t)-f^*\\le\\langle\\nabla f(w_t),w_t-w^*\\rangle =\\frac{R_t^2-R_{t+1}^2}{2\\eta}+\\frac\\eta2\\|\\nabla f(w_t)\\|^2. $$ Adding the descent lemma, the squared gradient term cancels when $\\eta=1/L$:\n$$ f(w_{t+1})-f^*\\le\\frac L2(R_t^2-R_{t+1}^2). $$ Summing from $t=0$ to $T-1$ yields $\\sum_{t=0}^{T-1}(f(w_{t+1})-f^{\\ast})\\le LR_0^2/2$. Since the function value is monotonically non-increasing under this step size, the left-hand side is at least $T(f(w_T)-f^{\\ast})$, proving the claim. The iteration complexity is $O(LR_0^2/\\epsilon)$, which is sublinear convergence.\n0x04 Non-smooth Convex with Subgradient Method Assume $f$ is convex, using\n$$ w_{t+1}=\\Pi_{\\mathcal C}(w_t-\\eta_tg_t),\\qquad g_t\\in\\partial f(w_t),\\quad\\|g_t\\|\\le G. $$ $\\Pi_{\\mathcal C}$ denotes the Euclidean projection; in the unconstrained case, simply omit the projection. We only require that the chosen subgradient is bounded; do not mistakenly call this condition smoothness.\nFixed Step Size for a Given Horizon Let $R=\\Vert w_0-w^{\\ast}\\Vert\u0026gt;0$. Given a priori $T\\ge1$, take a fixed step size $\\eta=R/(G\\sqrt T)$, then\n$$ \\min_{0\\le t\u0026lt;T}(f(w_t)-f^*)\\le\\frac{GR}{\\sqrt T}. $$ If only an upper bound on the initial distance is known, use that bound to set the step size; if the starting point is already optimal, stop immediately.\nProjection does not increase the distance to $w^{\\ast}\\in\\mathcal C$, hence\n$$ R_{t+1}^2\\le R_t^2-2\\eta(f(w_t)-f^*)+\\eta^2G^2. $$ Summing and dividing by $2\\eta T$:\n$$ \\min_{0\\le t\u0026lt;T}(f(w_t)-f^*) \\le\\frac{R^2}{2\\eta T}+\\frac{\\eta G^2}{2} =\\frac{GR}{\\sqrt T}. $$ Therefore, $T\\ge G^2R^2/\\epsilon^2$ is sufficient. This conclusion constrains the best iterate; by Jensen\u0026rsquo;s inequality, the unweighted average point also satisfies the same function value bound, but the last point is not guaranteed to have the same bound.\nPolyak Step Size If $f^{\\ast}$ is known and optimality has not yet been reached, one can take\n$$ \\eta_t=\\frac{f(w_t)-f^*}{\\|g_t\\|^2}. $$ Convexity provides the inequality $\\langle g_t,w_t-w^{\\ast}\\rangle\\ge f(w_t)-f^{\\ast}$. Substituting into the distance recurrence yields\n$$ R_{t+1}^2\\le R_t^2-\\frac{(f(w_t)-f^*)^2}{\\|g_t\\|^2}. $$ Summing and using $\\Vert g_t\\Vert\\le G$ gives $\\sum_{t=0}^{T-1}(f(w_t)-f^{\\ast})^2\\le G^2R^2$, so the best function value error is also bounded by $GR/\\sqrt T$. If $g_t=0$, convexity implies this point is already optimal; in this case, do not compute a step size with a zero denominator. In practice, $f^{\\ast}$ is often unknown, so this is a step-size rule requiring additional information.\n0x05 Smooth Non-convex with GD Assume $f$ is $L$-smooth and has a finite lower bound $f_{\\inf}$. Take $\\eta=1/L$; the descent lemma gives\n$$ f(w_{t+1})\\le f(w_t)-\\frac1{2L}\\|\\nabla f(w_t)\\|^2. $$ Summing from $t=0$ to $T-1$:\n$$ \\frac1T\\sum_{t=0}^{T-1}\\|\\nabla f(w_t)\\|^2 \\le\\frac{2L(f(w_0)-f_{\\inf})}{T}. $$ The best squared gradient also does not exceed this bound. Achieving $\\Vert\\nabla f\\Vert^2\\le\\epsilon$ requires $O(1/\\epsilon)$ iterations; if the target is $\\Vert\\nabla f\\Vert\\le\\epsilon$, then $O(1/\\epsilon^2)$. This is not a guarantee on the error of the global optimum value.\n0x06 Strongly Convex and Non-smooth with Subgradient Method Let $f$ be $\\mu$-strongly convex on the closed convex feasible set $\\mathcal C$, with the optimal solution attained in $\\mathcal C$. Along the projected subgradient iteration, we have $\\Vert g_t\\Vert\\le G$. Here, we do not require strong convexity and function value Lipschitz continuity simultaneously over the entire space.\nTake $\\eta_t=2/(\\mu(t+2))$; for $T\\ge1$, we have\n$$ \\min_{0\\le t\u0026lt;T}(f(w_t)-f^*)\\le\\frac{2G^2}{\\mu(T+1)}. $$ Proof: The subgradient inequality for strong convexity combined with projection yields\n$$ R_{t+1}^2\\le(1-\\mu\\eta_t)R_t^2-2\\eta_t(f(w_t)-f^*)+\\eta_t^2G^2. $$ Substitute the step size and rearrange:\n$$ f(w_t)-f^*\\le\\frac\\mu4\\left[tR_t^2-(t+2)R_{t+1}^2\\right]+\\frac{G^2}{\\mu(t+2)}. $$ Multiply by $t+1$ and sum; the distance terms telescope away, yielding\n$$ \\begin{aligned} \\sum_{t=0}^{T-1}(t+1)(f(w_t)-f^*) \u0026amp;\\le-\\frac\\mu4T(T+1)R_T^2 +\\frac{G^2}{\\mu}\\sum_{t=0}^{T-1}\\frac{t+1}{t+2}\\\\ \u0026amp;\\le\\frac{TG^2}{\\mu}. \\end{aligned} $$ Dividing by $\\sum_{t=0}^{T-1}(t+1)=T(T+1)/2$ gives the conclusion. The average point with equal weights also satisfies this function value bound. Its rate is $O(1/T)$, still sublinear.\n0x07 Conclusion Condition and Method Metric Guaranteed in This Paper Order of Error Upper Bound Smooth strongly convex, GD Function value error at the last iterate $O((1-\\mu/L)^T)$ Smooth convex, GD Function value error at the last iterate $O(1/T)$ Nonsmooth convex, subgradient method Function value error at the best point or average point $O(1/\\sqrt T)$ Smooth nonconvex, GD Average squared gradient $O(1/T)$ Nonsmooth strongly convex, projected subgradient method Function value error at the best point or weighted average point $O(1/T)$ These conclusions rely separately on the step sizes, lower bounds, existence of optimal solutions, or subgradient bounds discussed above; they cannot be applied merely by following the \u0026lsquo;convex/nonconvex\u0026rsquo; labels. General nonsmooth nonconvex problems require alternative algorithms and concepts of stationary points.\nReference Sources: 2 3 4 5.\nPart 1\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBoyd, S., Subgradient Methods.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBoyd, S. and Vandenberghe, L., Convex Optimization, Chapter 9.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBeck, A., First-Order Methods in Optimization.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNesterov, Y., Introductory Lectures on Convex Optimization: A Basic Course.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-convergence-analysis-in-deep-learning-part-2/","summary":"\u003ch2 id=\"0x00-preface\"\u003e0x00 Preface\u003c/h2\u003e\n\u003cp\u003eThis post analyzes deterministic first-order optimization methods. Basic definitions are in Part 1\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e. We default to the Euclidean norm and distinguish between gradient descent and the subgradient method for non-smooth problems.\u003c/p\u003e\n\u003cp\u003eThe smooth section discusses unconstrained problems; the non-smooth section allows a closed convex feasible set $\\mathcal C$, using projections to keep iterates within the feasible set. Assume an optimal solution $w^{\\ast}$ exists, $f^{\\ast}=f(w^{\\ast})$.\u003c/p\u003e\n\u003ch2 id=\"0x01-definition-of-convergence\"\u003e0x01 Definition of Convergence\u003c/h2\u003e\n\u003cp\u003eFirst, we must specify what error we are measuring:\u003c/p\u003e","title":"Convergence Analysis in Deep Learning (Part 2)"},{"content":"0x00 Preface This series organizes the convergence analysis of optimization methods. This installment unifies notation, introduces convexity, strong convexity, smoothness, stochastic gradient assumptions, and inequalities frequently used in subsequent proofs. I am not formally trained in optimization; I learned these topics out of research necessity. Please point out any omissions.\nThroughout this article, the Euclidean norm $\\Vert\\cdot\\Vert_2$ is used by default, abbreviated as $\\Vert\\cdot\\Vert$. Unless otherwise specified, we discuss unconstrained problems $\\min_{w\\in\\mathbb R^d}f(w)$. $w^{\\ast}$ denotes the optimal solution that exists, $f^{\\ast}=f(w^{\\ast})$; when the minimum may not be attained, we use a finite lower bound $f_{\\inf}$.\nThe objective functions of deep neural networks are typically non-convex. Theorems under convex or strongly convex conditions serve as theoretical models for understanding optimization and cannot be directly taken as guarantees for training general neural networks.\n0x01 Fundamental Concepts Convex Function A function $f$ is convex on a convex set $\\mathcal D$ if, for any $x,y\\in\\mathcal D$ and $\\theta\\in[0,1]$, we have\n$$ f(\\theta x+(1-\\theta)y)\\le\\theta f(x)+(1-\\theta)f(y). $$ When $f$ is differentiable, the equivalent first-order condition is\n$$ f(y)\\ge f(x)+\\langle\\nabla f(x),y-x\\rangle. $$ Non-smooth convex functions do not necessarily have gradients. In such cases, we use subgradients $g\\in\\partial f(x)$, which satisfy $f(y)\\ge f(x)+\\langle g,y-x\\rangle$ for all $y$.\n$\\mu$-strongly Convex defination\nFor a differentiable function, $\\mu$-strong convexity ($\\mu\u003e0$) is defined as $$ f(y)\\ge f(x)+\\langle\\nabla f(x),y-x\\rangle+\\frac{\\mu}{2}\\|y-x\\|^2. $$ In the non-smooth case, replace $\\nabla f(x)$ with $g\\in\\partial f(x)$. Under the Euclidean norm, $f$ being $\\mu$-strongly convex is equivalent to $f(x)-\\mu\\Vert x\\Vert^2/2$ being a convex function. We do not extend this equivalence to arbitrary norms here.\nIf $w^{\\ast}$ is the unconstrained optimal solution of a differentiable function, then $\\nabla f(w^{\\ast})=0$. Applying the strong convexity inequality separately yields\n$$ f(w)-f^*\\ge\\frac{\\mu}{2}\\|w-w^*\\|^2, \\qquad \\langle\\nabla f(w),w-w^*\\rangle \\ge f(w)-f^*+\\frac{\\mu}{2}\\|w-w^*\\|^2. $$ Lipschitz Continuous A function value $G$ being Lipschitz continuous means\n$$ |f(x)-f(y)|\\le G\\|x-y\\|. $$ On the entire space, for differentiable functions, this is equivalent to the gradient norm being uniformly bounded $\\Vert\\nabla f(x)\\Vert\\le G$. For non-smooth convex functions, the corresponding condition is bounded subgradients; on a restricted domain, one must specify which subgradients are selected and the boundary conditions.\nSmoothness defination\nA differentiable function $f$ is $L$-smooth, meaning its gradient $L$ is Lipschitz continuous: $$ \\|\\nabla f(x)-\\nabla f(y)\\|\\le L\\|x-y\\|. $$ Here, the constraint is on the variation of the gradient, which is a distinct property from Lipschitz continuity of the function values. Smoothness implies the descent lemma:\n$$ \\left|f(y)-f(x)-\\langle\\nabla f(x),y-x\\rangle\\right| \\le\\frac{L}{2}\\|y-x\\|^2. $$ The proof follows from line integration:\n$$ \\begin{aligned} f(y)-f(x)-\\langle\\nabla f(x),y-x\\rangle \u0026amp;=\\int_0^1\\langle\\nabla f(x+s(y-x))-\\nabla f(x),y-x\\rangle\\,ds,\\\\ \\left|f(y)-f(x)-\\langle\\nabla f(x),y-x\\rangle\\right| \u0026amp;\\le\\int_0^1 Ls\\|y-x\\|^2\\,ds =\\frac L2\\|y-x\\|^2. \\end{aligned} $$ For differentiable convex functions, a one-sided quadratic upper bound is equivalent to gradient Lipschitz continuity. However, for general non-convex functions, a one-sided upper bound cannot serve as this equivalence. For example, $f(x)=-x^4$ is a concave function that satisfies a one-sided upper bound for any positive $L$, yet its gradient is not Lipschitz over the entire space. Subsequent non-convex analysis adopts the gradient definition provided above.\n0x02 Optimization Methods Gradient Descent $$ w_{t+1}=w_t-\\eta_t\\nabla f(w_t),\\qquad\\eta_t\u0026gt;0. $$ Step sizes can be fixed or varying; the definition of gradient descent itself does not require the step size to be monotonically decreasing. The following identity serves as the starting point for distance analysis:\n$$ \\|w_{t+1}-w^*\\|^2 =\\|w_t-w^*\\|^2-2\\eta_t\\langle\\nabla f(w_t),w_t-w^*\\rangle +\\eta_t^2\\|\\nabla f(w_t)\\|^2. $$ Stochastic and Mini-batch Gradient Descent Let the loss for a single sample be $\\ell(w;\\xi)$, and the overall objective be $f(w)=\\mathbb E_\\xi[\\ell(w;\\xi)]$. Under conditions allowing the interchange of differentiation and expectation, independent sampling yields\n$$ g_t=\\frac1b\\sum_{r=1}^b\\nabla\\ell(w_t;\\xi_{t,r}),\\qquad w_{t+1}=w_t-\\eta_tg_t. $$ $b=1$ corresponds to single-sample SGD; $b\u0026gt;1$ corresponds to mini-batch SGD. We use $\\ell$ to distinguish sample losses from the objective function, avoiding confusion between single-sample gradients and batch averages.\n0x03 Common Assumptions These conditions are used as needed by specific theorems; not all analyses require them to hold simultaneously.\nBounded Domain and Bounded Subgradients If constrained optimization is required, one may assume the diameter of the closed convex feasible set $\\mathcal C$ does not exceed $D$, i.e., $\\Vert x-y\\Vert\\le D$. This is bounded domain, not bounded variance.\nNon-smooth analysis often assumes the selected subgradients satisfy $\\Vert g_t\\Vert\\le G$. A strongly convex function on the entire space cannot simultaneously have globally bounded gradients: strong convexity implies at least quadratic growth, while function value Lipschitz continuity allows at most linear growth. Therefore, relevant theorems must restrict the feasible set or the iteration trajectory and specify how constraints are maintained.\nConditional Unbiasedness and Bounded Variance Let $\\mathcal F_t$ denote the historical information before sampling, and $w_t$ be measurable with respect to $\\mathcal F_t$. Assume that new samples in each round are independent given the history, and\n$$ \\mathbb E[\\nabla\\ell(w_t;\\xi_{t,r})\\mid\\mathcal F_t]=\\nabla f(w_t), \\qquad \\mathbb E[\\|\\nabla\\ell(w_t;\\xi_{t,r})-\\nabla f(w_t)\\|^2\\mid\\mathcal F_t]\\le\\sigma^2. $$ Then the mini-batch average satisfies\n$$ \\mathbb E[g_t\\mid\\mathcal F_t]=\\nabla f(w_t),\\qquad \\mathbb E[\\|g_t-\\nabla f(w_t)\\|^2\\mid\\mathcal F_t]\\le\\frac{\\sigma^2}{b}, $$ $$ \\mathbb E[\\|g_t\\|^2\\mid\\mathcal F_t] \\le\\|\\nabla f(w_t)\\|^2+\\frac{\\sigma^2}{b}. $$ Variance reduction follows from the conditional expectation of the noise cross-terms in the average being zero. If samples are repeated or noise is correlated, unbiasedness alone does not imply that variance decreases by a factor of $1/b$. Here, vector variance is represented by the expected squared norm of the centered noise.\n0x04 Common Inequalities Cauchy-Schwarz and Triangle Inequalities $$ |\\langle x,y\\rangle|\\le\\|x\\|\\|y\\|,\\qquad \\|x+y\\|\\le\\|x\\|+\\|y\\|. $$ Jensen Inequality If $f$ is a convex function, $\\theta_i\\ge0$ and $\\sum_i\\theta_i=1$, then\n$$ f\\left(\\sum_i\\theta_ix_i\\right)\\le\\sum_i\\theta_if(x_i). $$ In the equal-weight case, it is $f(\\bar x)\\le n^{-1}\\sum_i f(x_i)$, where $\\bar x=n^{-1}\\sum_i x_i$. This does not imply $f(n\\bar x)\\le nf(\\bar x)$; one can construct a counterexample by taking $f(x)=x^2$.\nPolyak–Łojasiewicz (PL) Inequality defination\nThe PL condition is $$ \\|\\nabla f(x)\\|^2\\ge2\\mu(f(x)-f^*). $$ Differentiable, unconstrained $\\mu$-strongly convex functions satisfy this condition; the PL condition itself does not require the function to be strongly convex, nor does it require convexity. Proof that strong convexity implies PL: Fix $x$. Strong convexity yields, for any $y$,\n$$ f(y)\\ge f(x)+\\langle\\nabla f(x),y-x\\rangle+\\frac\\mu2\\|y-x\\|^2. $$ The minimum value of the right-hand side with respect to $y$ is $f(x)-\\Vert\\nabla f(x)\\Vert^2/(2\\mu)$. Taking the infimum on both sides gives $f^{\\ast}\\ge f(x)-\\Vert\\nabla f(x)\\Vert^2/(2\\mu)$; rearranging terms yields the conclusion.\n0x05 Acknowledgement I have been studying optimization content intermittently for nearly a year, and I have taken quite a few detours along the way. In the learning process, I initially started by studying derivations from papers, but many definitions were difficult to understand or grasp, so I turned to specialized optimization textbooks. Perhaps my aptitude is not the best, but many of these books left me confused or were too far removed from my research direction, making me somewhat impatient, which is why my progress was intermittent. In the second half of 2022, I took the public elective course Professor Niulingfeng titled Practical Optimization Algorithms and Their Applications, which finally gave me a proper understanding and overview of the entire field. This year, reading related content in the field has become much easier. I would like to thank Professor Niulingfeng for offering this course, which allowed me to easily get started and understand optimization-related content without having to enroll in specialized courses like Theory and Methods of Optimization.\nReference Sources: 1 2 3.\nBoyd, S. and Vandenberghe, L., Convex Optimization, Chapters 3 and 9.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nBottou, L., Curtis, F. E. and Nocedal, J., Optimization Methods for Large-Scale Machine Learning, Section 4.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nNesterov, Y., Introductory Lectures on Convex Optimization: A Basic Course.\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-convergence-analysis-in-deep-learning-part-1/","summary":"\u003ch2 id=\"0x00-preface\"\u003e0x00 Preface\u003c/h2\u003e\n\u003cp\u003eThis series organizes the convergence analysis of optimization methods. This installment unifies notation, introduces convexity, strong convexity, smoothness, stochastic gradient assumptions, and inequalities frequently used in subsequent proofs. I am not formally trained in optimization; I learned these topics out of research necessity. Please point out any omissions.\u003c/p\u003e\n\u003cp\u003eThroughout this article, the Euclidean norm $\\Vert\\cdot\\Vert_2$ is used by default, abbreviated as $\\Vert\\cdot\\Vert$. Unless otherwise specified, we discuss unconstrained problems $\\min_{w\\in\\mathbb R^d}f(w)$. $w^{\\ast}$ denotes the optimal solution that exists, $f^{\\ast}=f(w^{\\ast})$; when the minimum may not be attained, we use a finite lower bound $f_{\\inf}$.\u003c/p\u003e","title":"Convergence Analysis in Deep Learning (Part 1)"},{"content":"0x00 Preface The story goes like this: I\u0026rsquo;m about to return to the institute. There are three dormitory options: Suzhou Street, Youth Apartment, and Ke Yi Zhao. Among them, Ke Yi Zhao is likely the worst. Suzhou Street has the best living conditions, but it\u0026rsquo;s a 6-person room with a long commute time, making Youth Apartment seem quite OK.\nSince the quotas for Youth Apartment and Suzhou Street are limited, they are often allocated via a lottery. The most traditional lottery method is grabbing red envelopes. The strategy adopted by my lab is \u0026lsquo;winner takes all,\u0026rsquo; meaning those who get the Top K of WeChat red envelopes get the better accommodation.\n0x01 Analysis Bi Dao analyzed WeChat red envelope grabbing back in 2020 BV1z7411e7qB1, concluding that everyone\u0026rsquo;s expected value is the same, but the later you grab, the more likely you are to get a \u0026lsquo;big red envelope,\u0026rsquo; and the variance increases.\nThis leads to the probability of becoming the \u0026lsquo;Luckiest King\u0026rsquo;:\nUnder this condition, we hope not necessarily to be the Luckiest King, but to be in the Top K, so this can be considered an incremental work based on Bi Dao\u0026rsquo;s research.\n0x02 Simulation From Bi Dao\u0026rsquo;s video, we can see that WeChat red envelope amounts are distributed in [0.01, 2 * average of remaining amount]. Therefore, I used ChatGPT to write a simulation program, fixed some bugs myself, and here we only calculate for up to 20 people and the Top 10, simulating 100,000 times.\nHowever, there are some pitfalls to note: first, the amount grabbed must be rounded to two decimal places; second, if it\u0026rsquo;s the last person, they must grab the exact remaining amount.\nThe code is as follows:\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 import random import matplotlib.pyplot as plt # 模拟参数 num_trials = 100000 def simulate_red_envelope(num_users): total_amount = 100.0 red_envelope = [0.0] * num_users for i in range(num_users): remaining_envelope = num_users - i remaining_amount = total_amount - sum(red_envelope) avg_amount = remaining_amount / remaining_envelope max_amount = min(avg_amount * 2, remaining_amount) amount = ( random.uniform(0.01, max_amount) if remaining_envelope \u0026gt; 1 else remaining_amount ) amount = round(amount * 100) / 100.0 red_envelope[i] = amount return red_envelope def calculate_topk_probability(num_users, num_trials): probabilities = [] for k in range(1, min(num_users, 10)): results = [0] * num_users for _ in range(num_trials): red_envelope = simulate_red_envelope(num_users) sorted_envelope = sorted( range(num_users), key=lambda x: red_envelope[x], reverse=True ) topk = sorted_envelope[:k] # 统计获得 topk 的概率 for i in range(k): results[topk[i]] += 1 probabilities.append([result / num_trials for result in results]) return probabilities def plot_probability(probabilities): num_users = len(probabilities[0]) if len(probabilities) \u0026gt; 0 else 0 k_values = list(range(1, min(num_users, 10))) # 绘制 k 个子图 plt.figure(figsize=(4 * len(k_values) + 4, 4)) for i, k in enumerate(k_values): ax = plt.subplot(1, len(k_values), i + 1) ax.plot(range(1, num_users + 1), probabilities[i], label=\u0026#34;Probability\u0026#34;) print(probabilities[i], k) ax.set_xlabel(\u0026#34;Rank\u0026#34;) ax.set_xlim(1, num_users) # 只显示整数坐标 ax.set_xticks(range(1, num_users + 1)) ax.set_ylabel(\u0026#34;Probability\u0026#34;) ax.set_ylim(0, 1) # Title ax.set_title(\u0026#34;Probability of Top {} Amounts\u0026#34;.format(k)) # plt.xlabel(\u0026#34;Rank\u0026#34;) # plt.xlim(1, num_users) # # 只显示整数坐标 # plt.xticks(range(1, num_users + 1)) # plt.ylabel(\u0026#34;Probability\u0026#34;) # plt.ylim(0, 1) # plt.title(\u0026#34;Probability of Top K Amounts\u0026#34;) # plt.legend() plt.savefig(\u0026#34;red_envelope_probability_N{}_K{}.png\u0026#34;.format(num_users, k)) for N in range(2, 21): # 模拟抢红包并计算概率 probabilities = calculate_topk_probability(num_users=N, num_trials=num_trials) # 绘制概率图表 plot_probability(probabilities) 0x03 Results \u0026amp; Conclusions The results obtained are quite extensive; I will only display a few characteristic ones ($N = 2,3,4,5,10,20$).\nN=2\nN=3\nN=4\nN=5\nN=10\nN=20\nFirst, Top 1 is essentially the Luckiest King, used to compare with Bi Dao\u0026rsquo;s results to verify the correctness of my findings.\nRegarding the Luckiest King, just as Bi Dao concluded, the more people there are, the higher the probability that the last two people will become the Luckiest King.\nRegarding Top K, as K increases, this curve gradually changes from a concave curve to a convex curve, until finally becoming a monotonically decreasing curve. This is what is meant by \u0026lsquo;variance increases.\u0026rsquo; As K increases, the probability of the last person getting a lower amount increases, thus the probability of entering the Top K decreases. At the same time, the probability of entering the Top K also increases in mean value as K grows, after all, the probability of 10 people getting Top 10 is always 1.\nIn other words, within a certain range, it is better to grab later to get the Top K, but when K exceeds a certain value, it is better to grab earlier.\nSo, where does this threshold lie? Let\u0026rsquo;s explore this question next.\nThere are actually some small tricks here. First, the person who grabs last is very special because their amount is not obtained through sampling. When calculating the turning point, if we consider making the entire sequence monotonically decreasing, the sequence becomes extremely unstable, even lacking clear patterns (non-increasing), requiring more theoretical calculation to support this conclusion. It is also heavily influenced by sampling errors, as the probability difference between the last two is not significant. Limited by my knowledge of probability theory, I leave this difficult problem for the reader to ponder. However, if we do not consider the last person and only consider the decreasing sequence of the preceding ones, then this threshold becomes increasing as N increases.\nThe plot of N regarding the threshold k is shown below:\nCurve of the threshold k regarding N\nA conjectured conclusion is that this threshold k satisfies the following formula:\n$$ k = \\left \\lfloor \\frac{N-1}{4} \\right \\rfloor $$\nAs for why it is 4, it should be related to the distribution of WeChat red envelope amounts, but I lack theoretical analysis here.\n0x04 Limitations This is actually a setting of opposition among everyone. During the red envelope grabbing process, there is no information sharing, but in a real environment, you can ask classmates who have already grabbed to get the current number of people and the remaining amount. Under such conditions, the decision-making becomes more complex. For example, what should the decision be if previous people grabbed small red envelopes? What should the decision be if someone grabbed a very large red envelope? This problem remains to be explored and is left for the reader to think about.\nThere are still many points that have not been thoroughly studied, and limited by my probability knowledge, it is difficult to provide more probabilistic theoretical calculations. If readers are interested, welcome to discuss and exchange ideas in the comments section.\nBV1z7411e7qB\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-the-top-k-of-wechat-red-envelopes/","summary":"\u003ch2 id=\"0x00-preface\"\u003e0x00 Preface\u003c/h2\u003e\n\u003cp\u003eThe story goes like this: I\u0026rsquo;m about to return to the institute. There are three dormitory options: Suzhou Street, Youth Apartment, and Ke Yi Zhao. Among them, Ke Yi Zhao is likely the worst. Suzhou Street has the best living conditions, but it\u0026rsquo;s a 6-person room with a long commute time, making Youth Apartment seem quite OK.\u003c/p\u003e\n\u003cp\u003eSince the quotas for Youth Apartment and Suzhou Street are limited, they are often allocated via a lottery. The most traditional lottery method is grabbing red envelopes. The strategy adopted by my lab is \u0026lsquo;winner takes all,\u0026rsquo; meaning those who get the Top K of WeChat red envelopes get the better accommodation.\u003c/p\u003e","title":"How to Get the Top K of WeChat Red Envelopes"},{"content":"A Note Before We Begin This article primarily introduces methods for managing references and notes using Zotero and Notion.\nI am a heavy Notion user; I handle my daily notes and schedule management entirely within Notion. Plus, Notion offers an education discount, allowing you to use it for free indefinitely, so you don\u0026rsquo;t need to worry too much about storage space. Also, I quite like how Notion handles images and formulas.\nFollowing @YuYu\u0026rsquo;s recommendation, I tried using Zotero to manage my references. Discovering Notero for linking the two was just what I needed!\nStep by step First, create an integration. The name and logo can be anything you like. Open https://www.notion.so/my-integrations，点击 Create\nUnder Secrets, you will receive a token. Be careful not to leak this token. Manage permissions Duplicate the reference template library (click Duplicate in the top right corner). My template has been modified, so I am not using this exact one. Readers can modify it according to their needs: you can add items or delete Chinese tags, but do not change the English tags. https://slash-gem-e5d.notion.site/e629fd800387440b9e406b198b1520bd?v=f7e40abd4f2b4600a8510218efd9ed4f\nEstablish the connection Obtain the database ID. Click \u0026lsquo;Copy link\u0026rsquo; and save the database ID. note\nThe Notion link format is as follows: ```text https://notion.so/{workspace_name}/{database_id}?v={view_id} ``` Next, download Notero Open Notero\u0026rsquo;s release and download the latest version. Note that the file extension must be .xpi.\nInstall the plugin First, readers need to install Zotero and Zotero Connector. Then, after opening Zotero, go to\nTools - Add-ons - the small gear icon in the top right - Install Add-ons from file, and select the downloaded .xpi file.\nConfigure Notero Simply fill in the information from earlier.\nAll done~! Right-click a folder or a single item, and you will see \u0026lsquo;Sync Item to Notion\u0026rsquo;! It took me about half a day to complete the migration of my notes QAQ. I hope this will streamline my workflow in the future hhh\nFAQ I encountered one issue: when saving a reference, I got the error message \u0026lsquo;An error occurred while saving this item,\u0026rsquo; indicating a translator error. I downloaded the latest translator for Zotero, as well as some Chinese translator, but the problem persisted. Finally, I realized I had installed Zotero a long time ago and the version was far too outdated. Updating it solved the issue hhh\nReferences Sources: 1.\nhttps://zhuanlan.zhihu.com/p/455231476\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-using-zotero-and-notion-manage-your-papers-and-notes/","summary":"\u003ch2 id=\"a-note-before-we-begin\"\u003eA Note Before We Begin\u003c/h2\u003e\n\u003cp\u003eThis article primarily introduces methods for managing references and notes using Zotero and Notion.\u003c/p\u003e\n\u003cp\u003eI am a heavy Notion user; I handle my daily notes and schedule management entirely within Notion. Plus, Notion offers an education discount, allowing you to use it for free indefinitely, so you don\u0026rsquo;t need to worry too much about storage space. Also, I quite like how Notion handles images and formulas.\u003c/p\u003e","title":"Managing References and Notes with Zotero and Notion"},{"content":"Recently, something left a deep impression on me. It started when someone asked in a group chat if there was a good way to run multiple Python programs in parallel. However, most group members suggested using tmux, nohup, or simply appending \u0026amp; to run multiple processes directly. Listening to these suggestions, I felt none were as good as simply installing MPI, where a single line of code mpirun -np [num] python main.py could solve the problem directly. Professional parallel programs clearly offer better efficiency improvements and resource contention resolution compared to manually spawning multiple processes, and the number of processes can be controlled much better. Unfortunately, at least he himself likely did not adopt this solution. Even when I emphasized convenience, he seemed to have never considered it (of course, this is just my perspective; I can\u0026rsquo;t know if he searched for it himself, but he did not reply or ask further questions).\nThinking about it, this is quite normal. After all, for a completely unfamiliar program, when existing methods should suffice, how many people are willing to try something new? Including myself. I feel that as time passes and age grows, accepting a new concept might be easy, but the motivation to try a new product or software has declined significantly. Within my own knowledge domain, I am certain that the tool stack I already possess is sufficient to solve current problems. Unless someone specifically points it out or requires me to use it, the motivation to actually use it is not strong enough.\nRecently, while reading Sapiens, I had the same feeling. Perhaps future generations will look at us with the same sentiment: that they and \u0026lsquo;we\u0026rsquo; are all trapped within our own cognition. This viewpoint seems to be frequently mentioned. To break out of it, one must spend time trying and making mistakes, while maintaining a skeptical attitude towards problems (Are there better methods? Optimal results?).\nBut in today\u0026rsquo;s fast-paced era, can we really achieve this? Humans have constantly been \u0026lsquo;forcing\u0026rsquo; themselves and the world to speed up. Humans evolved from crawling to walking upright; the narrowing of female pelvises brought great pressure to childbirth, shortening gestation periods and simplifying fetal development in the womb. During the migration of Homo sapiens, whenever they arrived at a location, it inevitably triggered a wave of species extinction, especially for species with long reproductive cycles, as their breeding speed could not keep up with human hunting speeds. The agricultural era changed humans from infrequent hunting to daily farming labor, actually increasing the frequency of labor. The industrial era was no different. The arrival of the tech era is even more obvious: letters that once required careful word choice became instant messages available anytime. While work efficiency improved, it also meant an acceleration of work frequency. Do we really have enough time to experiment and make mistakes? Recently, I read several Q\u0026amp;A posts on Zhihu with similar sentiments. Those who truly settle down to do research from the basics often achieve little, whereas those who chase trends and do one thing at a time see results everywhere. I\u0026rsquo;m not saying either side is wrong or anything like that, but it feels that while both have their unique merits, the former often doesn\u0026rsquo;t get the results they deserve. As for myself, I always want to supplement courses and knowledge from the basics, but this is actually impossible. After all, human memory fades; it\u0026rsquo;s not like a course or something that requires frequent review. It\u0026rsquo;s easy to read something, get busy for a few days, and when you want to continue, you\u0026rsquo;ve forgotten almost everything. Now, I just use what I learn immediately. Furthermore, the cost of learning and accepting something new is clearly higher than just cobbling together a solution with existing viable methods. The deadline is there, the unknown is there; why bother? Additionally, there are simply too many new things. Take current AIGC products, for example; the explosion and surge have already exceeded my threshold for trying them out. Unless a product is particularly distinctive, at best, it just gathers dust in my bookmarks folder.\nActually, I have always questioned people of faith while also envying them; it\u0026rsquo;s quite contradictory. I believe that trusting a virtual, constructed deity, especially praying when facing difficulties, is something I view as a form of mental control, a means to reduce contradictions and suppress resistance. But conversely, this belief brings a delusion of escaping suffering and an unwillingness to give up. A person without faith might have a higher probability of giving up on their own life, after all, without the constraints of religious rules and ethics, a person as an independent individual might just \u0026lsquo;restart\u0026rsquo; if they can\u0026rsquo;t get over something. Yet, this very faith gave birth to science. This sounds absurd, as religion often played a role in the process of discovering science and correcting theories. The process of discovering science often requires skepticism first. Nowadays, this might be called being a \u0026lsquo;contrarian\u0026rsquo; or \u0026rsquo;troll\u0026rsquo;. Why does sometimes asking a rhetorical question bring such feedback? (Of course, here I mean genuine questioning, not sarcasm or passive-aggressive remarks.) To a certain extent, it\u0026rsquo;s because the other party cannot accept, is unwilling to accept, or simply doesn\u0026rsquo;t want to hear viewpoints outside their own cognition.\nIt seems that, besides our subjective unwillingness to accept, objective factors also constrain us, trapping us within our knowledge domains. In my view, there are two ways to break free: first, driven by incentives, such as the benefits from game trials or product trials, or if I try something out and can write a Zhihu article, WeChat official account post, or video tutorial to generate traffic and followers; or for research, if I read A and combine it with my field to publish an A+B paper. Second, driven by curiosity—trying everything I don\u0026rsquo;t understand. Personally, I feel this curiosity weakens with age, which is why I have a quote on my GitHub profile: \u0026lsquo;Stay curious :)\u0026rsquo;.\nHowever, being trapped in one\u0026rsquo;s own knowledge domain is not entirely a bad thing. If I use existing tools to create a new, more convenient tool, isn\u0026rsquo;t that also a transcendent way of thinking? Although the time cost might be higher, all roads lead to Rome; creating one\u0026rsquo;s own toolchain is not a bad thing.\n","permalink":"https://blog.bj-yan.top/en/p/misc-wo-men-zong-bei-kun-zai-zi-ji-de-ren-zhi-zhong/","summary":"\u003cp\u003eRecently, something left a deep impression on me. It started when someone asked in a group chat if there was a good way to run multiple Python programs in parallel. However, most group members suggested using tmux, nohup, or simply appending \u0026amp; to run multiple processes directly. Listening to these suggestions, I felt none were as good as simply installing MPI, where a single line of code \u003ccode\u003empirun -np [num] python main.py\u003c/code\u003e could solve the problem directly. Professional parallel programs clearly offer better efficiency improvements and resource contention resolution compared to manually spawning multiple processes, and the number of processes can be controlled much better. Unfortunately, at least he himself likely did not adopt this solution. Even when I emphasized convenience, he seemed to have never considered it (of course, this is just my perspective; I can\u0026rsquo;t know if he searched for it himself, but he did not reply or ask further questions).\u003c/p\u003e","title":"We Are Always Trapped in Our Own Cognition"},{"content":"Preface Recently, I was running experiments and found that the framework I had built previously was a bit uncomfortable to use, so I decided to refactor it. However, since the new framework isn\u0026rsquo;t ready yet, I first built a simple federated learning simulator based on mpi4py to run some basic experiments.\nThe code is hosted in this GitHub repository: https://github.com/yyyanbj/mpi-fedsim\nCode Structure The code structure is quite simple, mainly divided into five parts:\nmodel.py: Model definition, including the model\u0026rsquo;s definition and initialization utils.py: Some utility functions, as well as dataset downloading and splitting server.py and client.py: Definitions for the server and clients, including receiving and sending models, as well as model updates main_sync.py: The main program for synchronous federated learning algor: Some federated learning algorithms, including fedavg.py config: Some configuration and experimental parameters. The logging module uses loguru and tensorboard, while model creation and training are done with PyTorch. Although it\u0026rsquo;s quite simple, it still has \u0026ldquo;all the necessary organs,\u0026rdquo; hahaha.\nExecution Logic The logic mainly resides in the main function 1. First, if rank=0, it acts as the server; otherwise, it is a training process. The server initializes the model first, then waits for each process to upload the number of samples and the assigned client ID, recording them. Other processes first assign themselves a client ID based on pre-defined rules, then load the corresponding dataset based on that ID. Upon receiving the model, they begin training. After each process finishes training, it uploads the model to the server. Once the server receives it, it updates the global model and sends the model to all processes. Upon receiving the model, the processes update their local models and start the next round of training.\nSince my experiments are conducted in a single-machine multi-GPU environment, I can assign specific GPUs during thread allocation. However, since CNNs are quite small and won\u0026rsquo;t exhaust VRAM, I didn\u0026rsquo;t implement this part. If needed, you only need to modify the mapping relationship between main_sync.py and client_id with gpu_id.\nSince each process contains multiple clients but communicates only once when receiving model parameters, data transmission might be slightly faster.\nActually, the number of threads doesn\u0026rsquo;t really matter. The limit on threads stems from VRAM constraints, so as long as there is enough VRAM, the number of threads can be set very high. This approach simply aims to utilize as much VRAM as possible for acceleration within limited VRAM.\nUsage 1 2 3 4 5 6 7 8 9 10 11 12 13 14 # install requirements pip install -r requirements.txt # dataset prepare python utils.py # check and download the dataset you need # run the simulation # for linux mpirun -np 4 python main_sync.py # 4 is the number of processes # for windows mpiexec -n 4 python main_sync.py # 4 is the number of processes # launch the tensorboard tensorboard --logdir=logs Benchmark Here, a two-layer convolutional CNN is used for MNIST classification; I haven\u0026rsquo;t run too many tests. In practice, with 10 clients, each having 6,000 samples, a local epoch of 1, and 50 rounds, using 11 processes takes about 1 minute to complete.\nTo be added\nImprovements Actually, the code is still quite rough, and there are many parts that can be improved. First, regarding the model receiving and sending part, for synchronous learning, using a broadcast approach would be better, as it improves efficiency and eliminates the need for 1-to-1 checks.\nhttps://github.com/yyyanbj/mpi-fedsim/blob/main/main_sync.py\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-accelerating-federated-learning-simulation-using-mpi/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eRecently, I was running experiments and found that the framework I had built previously was a bit uncomfortable to use, so I decided to refactor it. However, since the new framework isn\u0026rsquo;t ready yet, I first built a simple federated learning simulator based on mpi4py to run some basic experiments.\u003c/p\u003e\n\u003cp\u003eThe code is hosted in this GitHub repository: \u003ca href=\"https://github.com/yyyanbj/mpi-fedsim\"\u003ehttps://github.com/yyyanbj/mpi-fedsim\u003c/a\u003e\u003c/p\u003e\n\u003ch2 id=\"code-structure\"\u003eCode Structure\u003c/h2\u003e\n\u003cp\u003eThe code structure is quite simple, mainly divided into five parts:\u003c/p\u003e","title":"Accelerating Federated Learning Simulation with MPI"},{"content":"Hey hey hey, after 600 hours in Apex Legends, I finally hit Diamond in solo queue.\nI started in S14 with a 0.6 KD. I reached Platinum in my very first ranked season, but after that, I mostly played casuals; ranked was basically \u0026lsquo;I give up\u0026rsquo; once I hit Platinum. This season, because the maps were rotating and reaching Platinum was relatively easy, I spent about a week trying to see if I could push to Diamond. It took around 50-60 hours, and boom, Diamond achieved 233. But I feel like this push to Diamond consumed all my enthusiasm, so now I\u0026rsquo;m trying to calm down - -.\nRank Up!\nLooking back at my years of gaming history: in elementary school, I played FPS games like CF and AVA; when everyone in middle school was playing LOL, I was still into FPS or Minecraft; I didn\u0026rsquo;t start playing LOL until high school; in university, I just switched between LOL and CSGO; and I only started playing Apex after seeing my roommates at UCAS play it. It seems I\u0026rsquo;m always one step late (except for Minecraft) 233.\nBut regarding ranked, it seems I\u0026rsquo;ve only reached Diamond twice: once in the PC version of Teamfight Tactics, and this time. Reaching Diamond in PC TFT also took a lot of grinding; I remember a Dawn Buff was too strong back then, getting strong synergies early and mid-game made me basically guaranteed top 4, and with a little late-game management and a formation switch, I had a chance to win. However, I\u0026rsquo;ve never reached Diamond in Summoner\u0026rsquo;s Rift 233. If Arena had ranked, I\u0026rsquo;m sure I could hit Diamond (confirmed).\nReaching Diamond in Apex solo queue is indeed not easy. I remember when I was Platinum I, I even got matched with Gold carrying Silver - -. My aim is just average; at least, I should have no problem one-tapping someone with a big side peek, big back peek, or finishing someone low on health. For close combat, being able to roam around is basically enough. My KD is also average, dropping from around 1.6+ when I hit Platinum to about 1.1 after reaching Diamond. The ranked maps this half-season were rotating; I played all three maps. Except for Crescent, where I got fewer wins, I got quite a few wins on the other two maps. I played different legends on different maps. On Crescent, the pressure to rotate into the circle was high, so I basically played Wattson or Gibraltar. Storm Point had too much verticality, and some circle shapes were easy to get stuck in, so I played Wattson more there. On World\u0026rsquo;s Edge, I basically played Gibraltar.\nIn solo queue, what I hope to see most is teammates who are willing to communicate and want to rank up. Some people like to roll for loot immediately upon landing. I don\u0026rsquo;t mind rolling with a team, but if you see 3 or 4 teams landing and you still insist on rolling, that basically means you don\u0026rsquo;t want to rank up. At that point, you need to remind them to rotate. You can try to persuade them; if they don\u0026rsquo;t rotate, just sell them and move on - -. Some teammates who play well will see other teams trying to rotate and run faster than me; it\u0026rsquo;s always good not to get caught by a rotation. I feel that in the latter half of my Platinum games, the quality of teammates was quite high; basically, everyone was willing to use voice chat. I would call out locations and damage for them. When pushing forward, I always called for my teammates, because sometimes they might not see the downed information.\nBesides communication, there\u0026rsquo;s also game sense. I learned some things about positioning, judging strong/weak sides, and understanding the circle mechanics from watching PUBG matches. Plus, since I mostly play Wattson and Gibraltar, who have good rotation capabilities, my decisions regarding entering the circle are generally correct. Some teammates can\u0026rsquo;t understand the pressure of rotating into a strong side; they push in against the pressure and basically either get stuck or get wiped out by rotations. I usually choose to lead teammates to the weak side and enter the circle first. If there\u0026rsquo;s nothing to do, just enter the circle to scout. If that doesn\u0026rsquo;t work, you can also choose not to rush into the circle immediately; just use the Bloodhound flow to push. This is basically learned from Apex pro matches (many people don\u0026rsquo;t actually understand this point), because the expected return from fighting early is just too low.\nSpeaking of expectations, you can calculate the expected risk and reward of fighting a team based on your ranked points; this basically tells you whether you should fight. If you have no resources, you must find a team to fight. It\u0026rsquo;s worth it to fight a team to enter the circle or to engage in a 3v3. If you are caught first and have to force a 3v3, you need to judge the situation.\nAlso, there\u0026rsquo;s the game environment. I feel like I occasionally encounter cheaters, but the probability is okay because basically cheaters have no brains; even if they are cheating, they can still get wiped by rotations. You can tell if someone is a cheater by their aim. A good player isn\u0026rsquo;t necessarily a cheater, but pre-firing is definitely a sign. You can also judge based on the distance and the damage dealt. If someone is cheating, just leave the game.\nAfter reaching Diamond, I haven\u0026rsquo;t dared to play Diamond ranked yet,after all, there are too many weird things - -. But after I finish my current work, I\u0026rsquo;ll try Diamond ranked. Maybe I\u0026rsquo;ll hit Master XD.\n","permalink":"https://blog.bj-yan.top/en/p/misc-games-rank-diamond/","summary":"\u003cp\u003eHey hey hey, after 600 hours in Apex Legends, I finally hit Diamond in solo queue.\u003c/p\u003e\n\u003cp\u003eI started in S14 with a 0.6 KD. I reached Platinum in my very first ranked season, but after that, I mostly played casuals; ranked was basically \u0026lsquo;I give up\u0026rsquo; once I hit Platinum. This season, because the maps were rotating and reaching Platinum was relatively easy, I spent about a week trying to see if I could push to Diamond. It took around 50-60 hours, and boom, Diamond achieved 233. But I feel like this push to Diamond consumed all my enthusiasm, so now I\u0026rsquo;m trying to calm down - -.\u003c/p\u003e","title":"Miscellaneous Notes: Diamond Rank"},{"content":"Introduction Today (2023.03.02), OpenAI released the latest GPT-3.5 Turbo API, currently priced at $0.002/1k tokens1.\nwarning\nThe information in this article is time-sensitive; please verify accordingly. Usage Guide The official documentation is here2, which is obviously more detailed than what I\u0026rsquo;ve written 233. Unless something unexpected happens, I still recommend checking the official docs.\nAfter writing this, I realized it might actually be more detailed than the official docs! 233\nRegistration warning\nYou may need a free internet environment for this. Also, if you are using proxy software, you must set `HTTP_PROXY` and `HTTPS_PROXY` in your command-line environment when running the program below, otherwise you will encounter access errors. On Windows, you can use the `set` command, `set HTTP_PROXY=http://127.0.0.1:xxxx`; on Linux, you can use the `export` command, `export HTTP_PROXY=http://127.0.0.1:xxxx`, where `xxxx` is the proxy software port number. First, register for a developer account on OpenAI Platform, then generate an API Key on the API Keys page.\nCreate API Key\nCurrently, OpenAI provides $18 in free credits for one month, which should be more than enough for testing.\nOne month of free credits\nDemo Example Let\u0026rsquo;s get started! Before beginning, you need to install the openai library; the current latest version is v0.27.0.\n1 pip install openai If you have installed it before, you might need to upgrade it.\n1 pip install openai --upgrade Below is a demo provided by the official documentation; simply replace the API key text with your own, and it will run directly.\n1 2 3 4 5 6 7 8 9 10 11 import os import openai os.environ[\u0026#34;OPENAI_API_KEY\u0026#34;] = \u0026#34;sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx\u0026#34; # replace with your API key openai.api_key = os.getenv(\u0026#34;OPENAI_API_KEY\u0026#34;) completion = openai.ChatCompletion.create( model=\u0026#34;gpt-3.5-turbo\u0026#34;, messages=[{\u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: \u0026#34;Hello!\u0026#34;}] ) print(completion.choices[0].message) Say Hi to GPT 3.5 Turbo\nThis way, we have successfully greeted GPT-3.5 Turbo!\nAdvanced Usage # 1 Let\u0026rsquo;s look at a longer example first.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 import os import openai os.environ[\u0026#34;OPENAI_API_KEY\u0026#34;] = \u0026#34;sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx\u0026#34; # replace with your API key openai.api_key = os.getenv(\u0026#34;OPENAI_API_KEY\u0026#34;) completion = openai.ChatCompletion.create( model=\u0026#34;gpt-3.5-turbo\u0026#34;, messages=[ {\u0026#34;role\u0026#34;: \u0026#34;system\u0026#34;, \u0026#34;content\u0026#34;: \u0026#34;You are a helpful assistant.\u0026#34;}, {\u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: \u0026#34;Who won the world series in 2020?\u0026#34;}, { \u0026#34;role\u0026#34;: \u0026#34;assistant\u0026#34;, \u0026#34;content\u0026#34;: \u0026#34;The Los Angeles Dodgers won the World Series in 2020.\u0026#34;, }, {\u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: \u0026#34;Where was it played?\u0026#34;}, ], ) print(completion.choices[0].message) First, you need to use the ChatCompletion from the openai library, then call the create method to create a ChatCompletion object, which contains our request information.\nLet\u0026rsquo;s observe the request format:\nmodel: No need to explain this; it is the model name. Currently, we are testing the gpt-3.5-turbo model, and only two models are supported: gpt-3.5-turbo and gpt-3.5-turbo-0301. Models with dates in their names will not be updated, but for today, the two are identical. messages: This is a list where each element is a dictionary. The role in the dictionary represents the message. Currently, three types are supported: user, system, and assistant. content represents the content of the message. system: System message, used to set the behavior of ChatGPT. user: User message, used to interact with ChatGPT. assistant: Assistant message, used to help store ChatGPT\u0026rsquo;s previous responses. Let\u0026rsquo;s look at the complete response.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 { \u0026#34;choices\u0026#34;: [ { \u0026#34;finish_reason\u0026#34;: \u0026#34;stop\u0026#34;, \u0026#34;index\u0026#34;: 0, \u0026#34;message\u0026#34;: { \u0026#34;content\u0026#34;: \u0026#34;The 2020 World Series was played at Globe Life Field in Arlington, Texas.\u0026#34;, \u0026#34;role\u0026#34;: \u0026#34;assistant\u0026#34; } } ], \u0026#34;created\u0026#34;: 1677759714, \u0026#34;id\u0026#34;: \u0026#34;chatcmpl-6pcEUG6zKP0Ld33OP1dZGbg3GwWhC\u0026#34;, \u0026#34;model\u0026#34;: \u0026#34;gpt-3.5-turbo-0301\u0026#34;, \u0026#34;object\u0026#34;: \u0026#34;chat.completion\u0026#34;, \u0026#34;usage\u0026#34;: { \u0026#34;completion_tokens\u0026#34;: 19, \u0026#34;prompt_tokens\u0026#34;: 56, \u0026#34;total_tokens\u0026#34;: 75 } } This is actually an example of multi-turn conversation. If you want to dynamically conduct multi-turn conversations, you must record and pass all previous responses each time.\nFirst, we set the system message, whose content is You are a helpful assistant.. Then we set the user message, whose content is Who won the world series in 2020?; this message serves as the information from the previous turn. Next, we set the assistant message, whose content is The Los Angeles Dodgers won the World Series in 2020., meaning ChatGPT\u0026rsquo;s response in the previous turn. Finally, we set the user message, whose content is Where was it played?, representing the question for this turn, i.e., the question we currently want GPT to answer.\nAs you can see, the received response content is The 2020 World Series was played at Globe Life Field in Arlington, Texas.. We only asked about the location, but ChatGPT already knew from the previous responses that this turn was about the 2020 World Series and answered accordingly.\nAt the same time, we can also see the usage field, which indicates how many tokens were used in this request: prompt_tokens is the number of input tokens, completion_tokens is the number of tokens in the ChatGPT response, and total_tokens is the total number of tokens used. This round consumed 75 tokens.\nAdvanced Usage # 2 Let\u0026rsquo;s look at a more complex example.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 import os import openai os.environ[\u0026#34;OPENAI_API_KEY\u0026#34;] = \u0026#34;sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx\u0026#34; # replace with your API key openai.api_key = os.getenv(\u0026#34;OPENAI_API_KEY\u0026#34;) completion = openai.ChatCompletion.create( model=\u0026#34;gpt-3.5-turbo\u0026#34;, messages=[ { \u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: \u0026#39;I want you to act as an Chinese translator, spelling corrector and improver. I will speak to you in any language and you will detect the language, translate it and answer in the corrected and improved version of my text, in Chinese. I want you to replace my simplified A0-level words and sentences with more beautiful and elegant, upper level Chinese words and sentences. Keep the meaning same, but make them more literary. I want you to only reply the correction, the improvements and nothing else, do not write explanations. My first sentence is \u0026#34;Non terrae plus ultra\u0026#34;\u0026#39;, }, ], temperature=0.9, # 0.0 to 2.0 (default 1.0) top_p=1, # 0.0 to 1.0 (default 1.0) (not used if temperature is set) n=5, # number (default 1) How many chat completion choices to generate for each input message. stream=False, # boolean (default False) stop=None, # string or array (default None) max_tokens=10, # inf (default 4096-prompt_token) presence_penalty=2.0, # -2.0 to 2.0 (default 0) frequency_penalty=0, # -2.0 to 2.0 (default 0) # logit_bias= # user= ) print(completion) for choice in completion.choices: print(choice.message.content) The prompt here is adapted from 3; the prompt itself is unrelated to the parameters I\u0026rsquo;m explaining.\nThis involves more parameters:\ntemperature: 0.0 to 2.0 (default 1.0) Temperature. Higher values make the output more random; lower values make it more deterministic (or regular). top_p: 0.0 to 1.0 (default 1.0) An alternative to temperature, also known as nucleus sampling. It is recommended not to use temperature and top_p simultaneously. top_p indicates that the model only considers the top top_p tokens by probability. For example, top_p=0.1 means the model only considers the top 10% of tokens by probability. n: number (default 1) The number of responses to generate. stream: boolean (default False) Whether to use streaming mode. If set to True, partial message chunks will be sent incrementally, just like in ChatGPT. What does that mean? It means a few words are sent to you one by one, allowing you to dynamically update the text as you wait for the full response in ChatGPT. stop: string or array (default None) Tokens used to stop generation. It can be a single string or a list of strings. If it\u0026rsquo;s a list, generation stops as soon as any one of the tokens appears, with a maximum of 4 tokens allowed. max_tokens: inf (default 4096 - prompt_token) The maximum number of tokens to generate. frequency_penalty and presence_penalty: -2.0 to 2.0 (default 0) Used to penalize repeated tokens. More details about these parameters are available in 4. One appears to handle frequency, while the other handles presence (as an integer). The higher the values of these two parameters, the less likely the generated text will repeat. The formula is as follows:\n1 mu[j] -\u0026gt; mu[j] - c[j] * alpha_frequency - float(c[j] \u0026gt; 0) * alpha_presence logit_bias: dict (default None) Used to adjust token probabilities; accepts JSON. Values range from -100 to 100. -100 effectively disables the word, while 100 forces its use if relevant. user: dict (default None) Used to set user information. For specifics, refer to 5, mainly to prevent abuse. The output of this code is as follows (since Chinese characters are escaped in JSON, I\u0026rsquo;ve replaced them with placeholders here).\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 { \u0026#34;choices\u0026#34;: [ { \u0026#34;finish_reason\u0026#34;: \u0026#34;stop\u0026#34;, \u0026#34;index\u0026#34;: 0, \u0026#34;message\u0026#34;: { \u0026#34;content\u0026#34;: \u0026#34;\\\u0026#34;天涯海角\\\u0026#34;\u0026#34;, \u0026#34;role\u0026#34;: \u0026#34;assistant\u0026#34; } }, { \u0026#34;finish_reason\u0026#34;: \u0026#34;stop\u0026#34;, \u0026#34;index\u0026#34;: 1, \u0026#34;message\u0026#34;: { \u0026#34;content\u0026#34;: \u0026#34;\\n\\n无穷尽之地\u0026#34;, \u0026#34;role\u0026#34;: \u0026#34;assistant\u0026#34; } }, { \u0026#34;finish_reason\u0026#34;: null, \u0026#34;index\u0026#34;: 2, \u0026#34;message\u0026#34;: { \u0026#34;content\u0026#34;: \u0026#34;\\n\\n\\\u0026#34;无出其右\\\u0026#34;\u0026#34;, \u0026#34;role\u0026#34;: \u0026#34;assistant\u0026#34; } }, { \u0026#34;finish_reason\u0026#34;: \u0026#34;stop\u0026#34;, \u0026#34;index\u0026#34;: 3, \u0026#34;message\u0026#34;: { \u0026#34;content\u0026#34;: \u0026#34;\\n\\n\\\u0026#34;无地不可至\\\u0026#34;\u0026#34;, \u0026#34;role\u0026#34;: \u0026#34;assistant\u0026#34; } }, { \u0026#34;finish_reason\u0026#34;: \u0026#34;length\u0026#34;, \u0026#34;index\u0026#34;: 4, \u0026#34;message\u0026#34;: { \u0026#34;content\u0026#34;: \u0026#34;\\n\\n无地可往，更远\u0026#34;, \u0026#34;role\u0026#34;: \u0026#34;assistant\u0026#34; } } ], \u0026#34;created\u0026#34;: 1677760749, \u0026#34;id\u0026#34;: \u0026#34;chatcmpl-6pcVBkXoD9xr2CSioR1Gz3ubEOdpg\u0026#34;, \u0026#34;model\u0026#34;: \u0026#34;gpt-3.5-turbo-0301\u0026#34;, \u0026#34;object\u0026#34;: \u0026#34;chat.completion\u0026#34;, \u0026#34;usage\u0026#34;: { \u0026#34;completion_tokens\u0026#34;: 45, \u0026#34;prompt_tokens\u0026#34;: 124, \u0026#34;total_tokens\u0026#34;: 169 } } Advanced Usage # 3 Here\u0026rsquo;s an example regarding the steam parameter.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 import os import openai os.environ[\u0026#34;OPENAI_API_KEY\u0026#34;] = \u0026#34;sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx\u0026#34; openai.api_key = os.getenv(\u0026#34;OPENAI_API_KEY\u0026#34;) completion = openai.ChatCompletion.create( model=\u0026#34;gpt-3.5-turbo\u0026#34;, messages=[ { \u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: \u0026#39;I want you to act as an Chinese translator, spelling corrector and improver. I will speak to you in any language and you will detect the language, translate it and answer in the corrected and improved version of my text, in Chinese. I want you to replace my simplified A0-level words and sentences with more beautiful and elegant, upper level Chinese words and sentences. Keep the meaning same, but make them more literary. I want you to only reply the correction, the improvements and nothing else, do not write explanations. My first sentence is \u0026#34;Non terrae plus ultra\u0026#34;\u0026#39;, }, ], temperature=1, # 0.0 to 2.0 (default 1.0) top_p=1, # 0.0 to 1.0 (default 1.0) (not used if temperature is set) n=1, # number (default 1) How many chat completion choices to generate for each input message. stream=True, # boolean (default False) stop=None, # string or array (default None) # max_tokens=100, # inf (default 4096-prompt_token) presence_penalty=2.0, # -2.0 to 2.0 (default 0) frequency_penalty=0, # -2.0 to 2.0 (default 0) # logit_bias= # user= ) for completion_ in completion: # print(completion_) for choice in completion_.choices: print(choice.delta.content if \u0026#34;content\u0026#34; in choice.delta else \u0026#34;\u0026#34;) When streaming mode is enabled, the response will be a stream of data rather than a single object containing all data, and this returned object is iterable. Below is an example of a returned item. Note that delta may not always contain content, so a check is necessary.\nIn other words, ChatGPT actually already has the complete result before outputting it; it just spits out words one by one to drag out the time on the frontend?!\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 { \u0026#34;choices\u0026#34;: [ { \u0026#34;delta\u0026#34;: { \u0026#34;content\u0026#34;: \u0026#34;\\uff1f\u0026#34; }, \u0026#34;finish_reason\u0026#34;: null, \u0026#34;index\u0026#34;: 0 } ], \u0026#34;created\u0026#34;: 1677763452, \u0026#34;id\u0026#34;: \u0026#34;chatcmpl-6pdCm5jwsB1e3YyEDZ1MQXbpHzWvn\u0026#34;, \u0026#34;model\u0026#34;: \u0026#34;gpt-3.5-turbo-0301\u0026#34;, \u0026#34;object\u0026#34;: \u0026#34;chat.completion.chunk\u0026#34; } Advanced Usage # 4 Here\u0026rsquo;s an example regarding the stop parameter.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 import os import openai os.environ[\u0026#34;OPENAI_API_KEY\u0026#34;] = \u0026#34;sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx\u0026#34; openai.api_key = os.getenv(\u0026#34;OPENAI_API_KEY\u0026#34;) completion = openai.ChatCompletion.create( model=\u0026#34;gpt-3.5-turbo\u0026#34;, messages=[ { \u0026#34;role\u0026#34;: \u0026#34;user\u0026#34;, \u0026#34;content\u0026#34;: \u0026#39;I want you to act as an Chinese translator, spelling corrector and improver. I will speak to you in any language and you will detect the language, translate it and answer in the corrected and improved version of my text, in Chinese. I want you to replace my simplified A0-level words and sentences with more beautiful and elegant, upper level Chinese words and sentences. Keep the meaning same, but make them more literary. I want you to only reply the correction, the improvements and nothing else, do not write explanations. My first sentence is \u0026#34;Non terrae plus ultra\u0026#34;\u0026#39;, }, ], temperature=1, # 0.0 to 2.0 (default 1.0) top_p=1, # 0.0 to 1.0 (default 1.0) (not used if temperature is set) n=1, # number (default 1) How many chat completion choices to generate for each input message. stream=False, # boolean (default False) stop=\u0026#34;无\u0026#34;, # string or array (default None) # max_tokens=100, # inf (default 4096-prompt_token) presence_penalty=0, # -2.0 to 2.0 (default 0) frequency_penalty=0, # -2.0 to 2.0 (default 0) # logit_bias= # user= ) for choice in completion.choices: print(choice.message.content) Here, we still have it play the role of a translator to translate a line from Apex Legends spoken by Octane. We use \u0026ldquo;无\u0026rdquo; (none) as the stop condition; when the output encounters \u0026ldquo;无\u0026rdquo;, generation stops and the result is returned. The output is as follows:\n1 您好，您所给出的第一句话“Non terrae plus ultra”是拉丁语，意为“没有比这更远的土地”，翻译成中文可写作“ As you can see, the output is directly cut off, and the interrupted word is automatically converted into a token. However, I haven\u0026rsquo;t yet thought of any practical application for this \u0026ndash;\nError I encountered this error several times while using it. It wasn\u0026rsquo;t due to issues with my request format; it was likely because too many people were making requests simultaneously. The error message looks like this:\n1 2 3 4 5 6 7 8 9 openai.error.APIError: The server had an error processing your request. Sorry about that! You can retry your request, or contact us through our help center at help.openai.com if you keep seeing this error. (Please include the request ID 7d8d2f6b67d92ff7850ef3e17d742827 in your email.) { \u0026#34;error\u0026#34;: { \u0026#34;message\u0026#34;: \u0026#34;The server had an error processing your request. Sorry about that! You can retry your request, or contact us through our help center at help.openai.com if you keep seeing this error. (Please include the request ID 7d8d2f6b67d92ff7850ef3e17d742827 in your email.)\u0026#34;, \u0026#34;type\u0026#34;: \u0026#34;server_error\u0026#34;, \u0026#34;param\u0026#34;: null, \u0026#34;code\u0026#34;: null } } 500 {\u0026#39;error\u0026#39;: {\u0026#39;message\u0026#39;: \u0026#39;The server had an error processing your request. Sorry about that! You can retry your request, or contact us through our help center at help.openai.com if you keep seeing this error. (Please include the request ID 7d8d2f6b67d92ff7850ef3e17d742827 in your email.)\u0026#39;, \u0026#39;type\u0026#39;: \u0026#39;server_error\u0026#39;, \u0026#39;param\u0026#39;: None, \u0026#39;code\u0026#39;: None}} {\u0026#39;Date\u0026#39;: \u0026#39;Thu, 02 Mar 2023 12:37:21 GMT\u0026#39;, \u0026#39;Content-Type\u0026#39;: \u0026#39;application/json\u0026#39;, \u0026#39;Content-Length\u0026#39;: \u0026#39;366\u0026#39;, \u0026#39;Connection\u0026#39;: \u0026#39;keep-alive\u0026#39;, \u0026#39;Access-Control-Allow-Origin\u0026#39;: \u0026#39;*\u0026#39;, \u0026#39;Openai-Model\u0026#39;: \u0026#39;gpt-3.5-turbo-0301\u0026#39;, \u0026#39;Openai-Organization\u0026#39;: \u0026#39;user-8qnumkqgd3l02hvzq5rqz0y1\u0026#39;, \u0026#39;Openai-Processing-Ms\u0026#39;: \u0026#39;750\u0026#39;, \u0026#39;Openai-Version\u0026#39;: \u0026#39;2020-10-01\u0026#39;, \u0026#39;Strict-Transport-Security\u0026#39;: \u0026#39;max-age=15724800; includeSubDomains\u0026#39;, \u0026#39;X-Request-Id\u0026#39;: \u0026#39;7d8d2f6b67d92ff7850ef3e17d742827\u0026#39;} Conclusion I can only say this pricing is incredibly cheap. It feels like many companies won\u0026rsquo;t even bother trying to replicate it themselves. Instead, they\u0026rsquo;ll just use the library. It offers good performance, no need to worry about costs, electricity, compute power, or other factors, and the price is low. I feel that even small companies with their own models might find that compute and electricity costs alone are significantly higher than the API. After all, this also comes down to utilization rates.\nAdditionally, I feel this has also changed the translation market to some extent. Taking Tencent Cloud\u0026rsquo;s translation API as an example, if you calculate the price roughly, it\u0026rsquo;s about 3 times that of the GPT-3.5 Turbo API. However, including input tokens, it\u0026rsquo;s more like 1.5 times, and it comes with other features, including rewriting and polishing.\nTencent Cloud Translation API pricing\nThrough experience, you can also find that if you input long texts every time, tokens are consumed quite quickly, as both input and output are billed. At the same time, if you want to conduct session-level conversations, your token consumption will grow rapidly: each turn multiplies by 2, and when accumulated, it becomes quadratic. So, the cost of long conversations is actually quite significant.\nHowever, unfortunately, OpenAI currently only supports virtual credit card payments. Domestic users who want to pay out of pocket may need to find their own solutions.\nDo you remember that last September I wrote a blog post about my thoughts on Stable Diffusion6? Now, looking at the SD models again, they seem to be a century behind\u0026hellip; LoRA, ControlNet\u0026hellip; If SD only impacted the art and design fields, then the potential of ChatGPT\u0026rsquo;s large models is truly vast and will affect many industries, because the diversity of outputs is incredibly rich. For example, someone might use it to generate code to drive machines, and so on7\u0026hellip; Application scenarios depend entirely on imagination. However, there are currently scientific issues. If a more powerful knowledge base could be established to provide theoretical backing during output, the application scenarios of this model would become even broader, such as in healthcare, finance, and other fields with higher demands for evidence and decision-making.\nhttps://openai.com/blog/introducing-chatgpt-and-whisper-apis\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://platform.openai.com/docs/guides/chat\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://github.com/f/awesome-chatgpt-prompts\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://platform.openai.com/docs/api-reference/parameter-details\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://platform.openai.com/docs/guides/safety-best-practices/end-user-ids\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://blog.bj-yan.top/p/blog-will-ai-replace-the-artists/\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://github.com/microsoft/PromptCraft-Robotics\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-advantage-usage-for-gpt-3-5-turbo/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eToday (2023.03.02), OpenAI released the latest GPT-3.5 Turbo API, currently priced at $0.002/1k tokens\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e.\u003c/p\u003e\n\n  \u003cblockquote class=\"book-hint2 warning\"\u003e\n    \u003cp class=\"hint-title warning\"\u003e\n      \u003csvg class=\"book-icon\"\u003e\n        \u003cuse href=\"/svg/hint-icons.svg#warning-notice\"\u003e\u003c/use\u003e\n      \u003c/svg\u003e\u003cspan\u003ewarning\u003c/span\u003e\u003c/p\u003e\n    \nThe information in this article is time-sensitive; please verify accordingly.\n\n  \u003c/blockquote\u003e\n\n\u003ch2 id=\"usage-guide\"\u003eUsage Guide\u003c/h2\u003e\n\u003cp\u003eThe official documentation is here\u003csup id=\"fnref:2\"\u003e\u003ca href=\"#fn:2\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e2\u003c/a\u003e\u003c/sup\u003e, which is obviously more detailed than what I\u0026rsquo;ve written 233. Unless something unexpected happens, I still recommend checking the official docs.\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eAfter writing this, I realized it might actually be more detailed than the official docs! 233\u003c/p\u003e","title":"Guide to Using the GPT-3.5 Turbo API: Parameters and Code Examples (2023)"},{"content":" Here come some useless knowledge! This article mainly records some tips on GitHub.\nWhy use GitHub? Code management Get GitHub Pages and Actions for free! Special Repositories There are mainly two types of special repositories: one is username.github.io, and the other is username/username (both personal and organization accounts have these two types of repositories).\nusername.github.io The main purpose of this type of repository is to host personal blogs, but it can also be used to host other things, such as personal websites, documentation for personal projects, etc. GitHub Pages will automatically host the content of this repository at the address https://username.github.io.\nusername/username This type of repository is generally used to host information about a personal account, such as a profile. Open your homepage; the content of the file README.md in this repository is your profile.\nReference: yyyanbj - GitHub1\nFor organization accounts, it is slightly different. You need to place it in the file profile/README.md within the repository organization/.github.\nReference: awesome-actions-template - GitHub2\nGitHub Pages Custom Domain Generally, you need to place a file named CNAME in the root directory of your repository\u0026rsquo;s GitHub Pages, write your domain name inside it, and then add a CNAME record at your domain\u0026rsquo;s DNS provider. For a subdomain like www, point it to username.github.io. Additionally, add a A record for a subdomain like @, pointing to the IP address of GitHub Pages. If you are using a blog engine, you should place the CNAME file in the static directory of your blog engine.\n1 2 3 4 185.199.108.153 185.199.109.153 185.199.110.153 185.199.111.153 warning\nNote that the IP addresses here are time-sensitive. If you find these IP addresses have expired, you can find the latest ones at Managing a custom domain for your GitHub Pages site - GitHub Docs[^cite-3]. If you are an IPv6 user, you also need to add a AAAA record. For a subdomain like @, point it to the IPv6 address provided in Managing a custom domain for your GitHub Pages site - GitHub Docs3.\nGitHub Actions github-actions[bot] Now that GitHub Actions has been updated, permission management has become more granular. By default, github-actions[bot] often does not have write permissions for the repository and needs to be added manually.\nAddition method: Settings -\u0026gt; Actions -\u0026gt; General -\u0026gt; Workflow permissions -\u0026gt; Read and write permissions -\u0026gt; Save\nPro Account If you are a student, you can apply for a free Pro account with a Pro Plan at education.github.com. You can then use GitHub Actions and GitHub Pages in the Private repository.\nGitHub Emoji GitHub has its own Emoji, which can be found at GitHub Emoji Cheat Sheet4 and used in issues, commit messages, Discussions, etc. Of course, you can also use them in the comment section below, which is powered by giscus!\nConclusion Updated irregularly QAQ\nyyyanbj - GitHub\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nawesome-actions-template - GitHub\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nManaging a custom domain for your GitHub Pages site - GitHub Docs\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGitHub Emoji Cheat Sheet\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-github-useless-tips/","summary":"\u003cblockquote\u003e\n\u003cp\u003eHere come some useless knowledge! This article mainly records some tips on GitHub.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"why-use-github\"\u003eWhy use GitHub?\u003c/h2\u003e\n\u003cul\u003e\n\u003cli\u003eCode management\u003c/li\u003e\n\u003cli\u003eGet GitHub Pages and Actions for free!\u003c/li\u003e\n\u003c/ul\u003e\n\u003ch2 id=\"special-repositories\"\u003eSpecial Repositories\u003c/h2\u003e\n\u003cp\u003eThere are mainly two types of special repositories: one is \u003ccode\u003eusername.github.io\u003c/code\u003e, and the other is \u003ccode\u003eusername/username\u003c/code\u003e (both personal and organization accounts have these two types of repositories).\u003c/p\u003e\n\u003ch3 id=\"usernamegithubio\"\u003e\u003ccode\u003eusername.github.io\u003c/code\u003e\u003c/h3\u003e\n\u003cp\u003eThe main purpose of this type of repository is to host personal blogs, but it can also be used to host other things, such as personal websites, documentation for personal projects, etc. GitHub Pages will automatically host the content of this repository at the address \u003ccode\u003ehttps://username.github.io\u003c/code\u003e.\u003c/p\u003e","title":"Useless Tips for GitHub"},{"content":" It\u0026rsquo;s been half a year, and I\u0026rsquo;ve finally updated this project again!!!\nWhen I first finished the initial version of this project, I wrote a related 1. Since then, I\u0026rsquo;ve made sporadic modifications, but most were only updated locally. This time, I made some changes that led to this v0.1.0 version.\nActually, I\u0026rsquo;ve used this library for many of my projects (after all, there\u0026rsquo;s no need to reinvent the wheel; it\u0026rsquo;s truly satisfying), such as fedhf. However, I often got annoyed by my own __frozen__ attribute, so I simply deleted it everywhere you can see (including output and comparison). The comparison logic has also been revised; you can now compare directly with dictionaries by ignoring this parameter.\nAdditionally, I fixed something that gave me a headache: the load method. I initially thought instantiating an object and then loading it would be fine, but after using it for a while, I found it was still troublesome. Every time, I had to import a Config object before loading, which felt redundant. So, I simply wrote a load function that can be called directly without instantiation.\nHere are the updates from the first version to now: v0.0.1...v0.1.0\nIn fact, after using it for a while, I realized that assignments like cfg.a.b.c = 1 are not very common for me; querying is usually the most frequent scenario.\nI should continue to update it irregularly in the future; most updates will be driven by my own needs. I should also find time to work on ez.save(). I just forgot and wrote it down on the spot QAQ. Also, I need to handle compatibility with pathlib, since pathlib\u0026rsquo;s Path objects are also commonly used by me.\nI happened to update the issues mentioned above, and now it\u0026rsquo;s not v0.1.0 anymore; it\u0026rsquo;s v0.1.1.\nHere are the most recent updates: v0.1.0...v0.1.1\nThe handling of pathlib should probably be moved to a separate utils module. Depending on my future usage, if it proves convenient and no other requirements arise, there likely won\u0026rsquo;t be major changes.\ninfo\nIf you have any good suggestions or ideas, feel free to leave a comment below or submit an issue directly on [GitHub](https://github.com/yyyanbj/ezkfg). blog\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-public-release-of-ezkfg-v011/","summary":"\u003cblockquote\u003e\n\u003cp\u003eIt\u0026rsquo;s been half a year, and I\u0026rsquo;ve finally updated this project again!!!\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eWhen I first finished the initial version of this project, I wrote a related \u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e. Since then, I\u0026rsquo;ve made sporadic modifications, but most were only updated locally. This time, I made some changes that led to this \u003ccode\u003ev0.1.0\u003c/code\u003e version.\u003c/p\u003e\n\u003cp\u003eActually, I\u0026rsquo;ve used this library for many of my projects (after all, there\u0026rsquo;s no need to reinvent the wheel; it\u0026rsquo;s truly satisfying), such as \u003ca href=\"https://github.com/yyyanbj/fedhf\"\u003e\u003ccode\u003efedhf\u003c/code\u003e\u003c/a\u003e. However, I often got annoyed by my own \u003ccode\u003e__frozen__\u003c/code\u003e attribute, so I simply deleted it everywhere you can see (including output and comparison). The comparison logic has also been revised; you can now compare directly with dictionaries by ignoring this parameter.\u003c/p\u003e","title":"ezkfg v0.1.1 Release: Configuration Loading and Saving Updates"},{"content":"Introduction If you work in the field of deep learning, you are likely not unfamiliar with the recently popular Diffusion models. On the popular machine learning model hosting website Hugging Face, the top entries in the Trending section are, without exception, all diffusion models.\nThe images generated by these diffusion models are equally stunning. Whether landscapes, portraits, or paintings in various styles, the generated images exhibit extremely high quality. Here are a few examples generated by these models.\nImage source: Midjourney Hanzo\nImage source: Midjourney LiaLöwenherz🦁💙\nOf course, many large and small enterprises have already begun utilizing diffusion models for commercial operations. For example, abroad:\nMidjourney Stability AI Domestically:\n6pen.art Baidu - Wenxin ERNIE-ViLG Some have even used AI to create images celebrating the LOL championship victory 1, and even People\u0026rsquo;s Daily released a Mid-Autumn Festival video2 generated using diffusion models on Bilibili, proving it has truly gone viral.\nRecently, while taking a course on Dialectics of Nature, I wanted to reflect on this issue, which led to this article. For a brief history, principles, and performance comparison of diffusion models, please refer to my other article3.\nThe Question Raised Given how stunning these models perform, might one day we no longer need artists to create art?\nWill the work of artists be replaced by such generative art? In 5 years? 10 years? 20 years? 100 years? What will future art and artists look like?\nThe Rapid Development of AI The history of AI dates back a long way. Here, I will briefly outline the significant advancements in the AI field, particularly deep learning, in recent years, along with several milestones:\n2012: The AlexNet model surpassed humans on the ImageNet image classification task. 2016: AlphaGo, in a field humans never believed a machine could defeat them—Go—soundly defeated world Go champion Lee Sedol. 2018: StyleGAN generated human faces indistinguishable from real ones to the human eye. 2021: OpenAI\u0026rsquo;s GPT-3, a large language generation model based on context, capable of generating various types of text, even code. It is said to have cost OpenAI $4.6 million to train 4. 2021: DeepMind\u0026rsquo;s AlphaFold 2, a protein structure prediction model with high accuracy for unknown proteins, aiding drug development. 2022: Stable Diffusion (aka DALL·E 2), high-quality text-to-image generation. AI has already defeated humans in many fields and is beginning to assume a position of dominance.\nThis CLIP + Diffusion model has also shattered many preconceived notions. Just as AlphaGo\u0026rsquo;s arrival did, before which people believed no model could defeat humans in a complex field like Go, the result is now well known: AlphaGo defeated world Go champion Lee Sedol 3-1. Since then, no one has been able to defeat AI in the field of Go.\nAfter AlphaGo defeated Lee Sedol, Ke Jie once complained during a live stream that AI has made Go very \u0026lsquo;boring\u0026rsquo; 5. Now, it\u0026rsquo;s not about how skilled you are, but how much you understand AI, how similar you can become to AI, and learning to play Go like AI. Yet, it cannot be denied that AI has also led to tremendous progress for humans in the field of Go.\nThis Diffusion model brings a shock similar to what AlphaGo did. Previously, many believed deep learning models would lack creativity, merely inducting and deducing existing knowledge, unable to create something never seen before (including myself). However, here, it not only understands descriptions in strange languages, no matter how fantastical, and generates reasonable images that fit the description, but it can also generate things never seen or events impossible in the human worldview, and even create things with great creativity and imagination.\nImage source: Midjourney Discord community BartonDH 6. This image was even taken to OpenSea to be sold as an NFT, eventually getting reported and taken down, highlighting that copyright issues in the NFT market remain extremely difficult to resolve.\nWhich Jobs Will Be Replaced From ancient times to the present, machines replacing manual labor to improve production efficiency is inevitable.\nIn January 2019, scholars Baobao Zhang and Allan Dafoe from the Center for the Governance of AI at the Oxford University Future of Humanity Institute released an 111-page report titled \u0026ldquo;Artificial Intelligence: American Attitudes and Trends\u0026rdquo; 7 8, which mentioned the risk of AI replacing certain repetitive jobs. In 2013, Frey et al.\u0026rsquo;s \u0026ldquo;The future of employment: How susceptible are jobs to computerisation?\u0026rdquo; 9 listed over 700 occupations and their probabilities of being replaced, noting that 47% of jobs in the United States faced a high risk of replacement, including telemarketers, title screeners, textile workers, and so on. Most of these are highly repetitive jobs; for instance, telemarketing largely involves making repetitive calls using nearly identical scripts. Nowadays, many of the nuisance calls we receive are no longer made by humans. In contrast, they considered creative professions such as artists and scientists to have a lower probability of being replaced.\nOf course, this article is from 2013, and many of its viewpoints are now somewhat outdated; professions previously deemed to have an extremely low risk of replacement now face significant risks. For example, the field of AI in Science has recently become extremely popular. Take the December 2021 article on the cover of Nature, \u0026ldquo;Advancing mathematics by guiding human intuition with AI\u0026rdquo; 10, which used AI to guide the proof of mathematical formulas. Although this does not mean AI can already replace mathematicians, it has begun guiding the mathematical intuition of scientists and mathematicians, possessing a certain level of mathematical literacy, and helping mathematicians gain inspiration for proving theorems. In the future, using AI to guide intuition and improve research efficiency is certainly not a pipe dream.\nHow far can AI go? What problems arise? So, returning to art, the core logic of current diffusion models remains largely unchanged: they can generate a complete image from random noise or an initial value.\nCavemen taking a group selfie\nAn astronaut, riding a horse, in a photorealistic style\nAnd beyond generation 11 12, there are already numerous painting skills, such as image inpainting, image super-resolution, and editing images based on text descriptions.\nThis is a completion of \u0026ldquo;Girl with a Pearl Earring\u0026rdquo; by OpenAI\u0026rsquo;s Dall·E 13.\nSuper-resolution from the Imagen paper 14.\nImage editing by Dall·E-2 12. The fact that AI can reach this level carries certain risks. After all, AI can understand the meanings of specific terminology because its training data mostly comes from internet images rather than fixed datasets. This brings up many copyright issues (beyond copyright, there are also rights of portrait, rights related to characters, and various other ownership rights), sparking significant online debate and leading to resistance against AI-generated art by many illustrators 15 16.\nGenerated by Stable Diffusion, prompt: trump kiss putin\nMany illustrators or concept artists\u0026rsquo; livelihoods depend on their unique artistic styles. However, a very realistic and easy thing is that after an artist spends years designing an elegant and refined style, AI can glance at it, train on the data, and effortlessly generate a pile of works in the same style within seconds. Is this creative efficiency somewhat unbalanced? It is only natural and understandable that artists would resist this 17. After all, such products are ultimately intended for commercial operation, while artists\u0026rsquo; works are released on the internet for free and absorbed and learned by AI, which is clearly unfair to the artists.\nDALL·E 2 VARIATIONS of The Girl with a Pearl Earring\nHowever, this is not up to the artists to decide; existing laws cannot prohibit such creation. After all, imitating a style does not count as plagiarism or infringement. At least under current Chinese law, there are no provisions protecting rights in this area. This also shows that \u0026ldquo;ethical construction lags far behind the pace of technological development.\u0026rdquo; Based on existing legal cases in China, there are two aspects: one is the affirmation of copyright for AI-generated works, and the other is the definition of \u0026ldquo;originality.\u0026rdquo;\nThe law recognizes and protects the copyright of AI-created works 18, affirming the work involved in AI model training and prompt tuning (often referred to as \u0026ldquo;alchemy\u0026rdquo;). In the process of AI generating artwork, it only uses data for training and incorporates certain elements of the original works into the final product, which meets the criteria of \u0026ldquo;originality,\u0026rdquo; just as humans inevitably produce works with similar styles after observing other pieces 19.\nRegarding international recognition of AI originality, a very typical example occurred recently. In August 2022, at an art fair in Colorado, USA, \u0026ldquo;Théâtre D\u0026rsquo;opéra Spatial\u0026rdquo; won the champion in the digital art category 20.\n《Théâtre D\u0026rsquo;opéra Spatial》 by Jason Allen via Midjourney\nAI will undoubtedly bring massive disruption to the field of painting, and it\u0026rsquo;s not just about that; there will also be significant impacts on the anime and film/TV industries. Although currently generating highly continuous video is not yet possible, methods like interpolation and frame filling exist. As research deepens, AI-generated video will be realized very soon. There are already demos of AI performing video editing 21 and AR 22.\nAnother point is the risk of other elements appearing in AI-generated images, such as pornography, gore, violence, etc. For example, Reddit has already banned many NSFW posts 23.\n\u0026ldquo;Creator\u0026rdquo; + Technical Skill = \u0026ldquo;Artist\u0026rdquo; So, while AI-generated images currently carry some risks, this is a rapidly iterating field, and these risks will gradually be mitigated in future developments. But it cannot be denied that many AI-generated images are already imaginative and creative. So, perhaps we truly no longer need \u0026ldquo;artists\u0026rdquo;?\nLet\u0026rsquo;s return to the most fundamental question: what is an artist, 将其所体验的世界通过各艺术种类的独特艺术语言和表现手段转化成艺术作品的人 is called an artist 24. In this process, as AI develops, art forms, artistic languages, and expressive techniques will inevitably occupy every domain of art, such as music, watercolor, oil painting, sketching, and so on. But what is most crucial is the ideas that artists derive from their experiences of the world. Regardless of the form, the form itself is merely one of the artist\u0026rsquo;s expressive techniques.\nTherefore, I believe that \u0026ldquo;creators\u0026rdquo; may eventually replace the status of \u0026ldquo;artists,\u0026rdquo; diminishing the emphasis on artistic skills while focusing on content expression.\nIn the short term, artists can use AI as a tool to quickly generate initial works for iterative upgrades and refinement 25. However, in the medium to long term, it is inevitable that AI will replace artists\u0026rsquo; work. It allows people without technical skills to express their thoughts and viewpoints through various forms, enabling more people to participate in artistic creation.\nCurrently, on Bilibili, there are already \u0026ldquo;creators\u0026rdquo; who use AI to generate images and videos for profit 26. At the same time, many derivative professions have emerged: for example, teaching prompt engineering, selling high-quality image prompts, etc. There are also prompt search websites like openart and prompthero , truly giving rise to the role of \u0026ldquo;prompt engineers\u0026rdquo;.\nConclusion Although AI cannot yet completely replace all of an \u0026ldquo;artist\u0026rsquo;s\u0026rdquo; work, it can already serve as a tool to provide more people with avenues for artistic expression, greatly enhancing the productivity of \u0026ldquo;creators.\u0026rdquo; In the near future, diffusion models will certainly offer more detailed optimizations, such as lighting, perspective, etc., to enable finer scene control. At that time, everyone can become a \u0026ldquo;creator,\u0026rdquo; displaying their inspiration to the public at any moment.\nOf course, at the same time, we should also focus on protecting artists\u0026rsquo; copyrights and ideas. We can leverage technologies and methods like NFTs and Web 3.0 to improve relevant laws and regulations, safeguarding artists\u0026rsquo; original works.\nHowever, in the distant future, if AI truly develops to the point where it possesses its own emotions and coexists with humans, it will certainly be able to replace all of an \u0026ldquo;artist\u0026rsquo;s\u0026rdquo; work. Then the singularity of artificial intelligence will have arrived 27. We can see that art and science are two great peaks; if AI conquers art, then ultimately surpassing humanity is not far off. Right now, we cannot deduce whether this will be a blessing or a disaster.\nReference https://lol.qq.com/news/detail.shtml?type=6\u0026docid=4934556576507716833\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nMid-Autumn Festival video\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nother article\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.sohu.com/a/429205048_120828615\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://view.inews.qq.com/a/20220514A087C500\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://discord.com/channels/662267976984297473/1008049088324972657/1015362328906182748\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://isps.yale.edu/sites/default/files/files/Zhang_us_public_opinion_report_jan_2019.pdf\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.zhihu.com/zvideo/1326127307812700160\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://ora.ox.ac.uk/objects/uuid:4ed9f1bd-27e9-4e30-997e-5fc8405b0491/download_file?safe_filename=future-of-employment.pdf\u0026file_format=application%2Fpdf\u0026type_of_work=journal+article\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.nature.com/articles/s41586-021-04086-x.pdf\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.reddit.com/r/midjourney/comments/wvoscd/cavemen_taking_a_group_selfie/\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://openai.com/dall-e-2/\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://openai.com/blog/dall-e-introducing-outpainting/\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://arxiv.org/abs/2205.11487\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.gamersky.com/ent/202208/1513964.shtml\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.zhihu.com/question/550660606\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.zhihu.com/question/550997249/answer/2656595328\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://zhuanlan.zhihu.com/p/565071999\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.zhihu.com/question/552231525/answer/2665147875\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://baike.baidu.com/item/%E5%A4%AA%E7%A9%BA%E6%AD%8C%E5%89%A7%E9%99%A2/61959625?fr=aladdin\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://twitter.com/runwayml/status/1568220303808991232\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://twitter.com/StrangeNative/status/1569700294673702912\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://zhuanlan.zhihu.com/p/560232893\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://baike.baidu.com/item/%E8%89%BA%E6%9C%AF%E5%AE%B6/23418?fr=aladdin\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://zhuanlan.zhihu.com/p/378444440\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://space.bilibili.com/335884771\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.zhihu.com/question/284243786/answer/1131987569\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-will-ai-replace-the-artists/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eIf you work in the field of deep learning, you are likely not unfamiliar with the recently popular Diffusion models. On the popular machine learning model hosting website \u003ca href=\"https://huggingface.co/\"\u003eHugging Face\u003c/a\u003e, the top entries in the Trending section are, without exception, all diffusion models.\u003c/p\u003e\n\n\u003cfigure class=\"align-center\"\u003e\u003ca href=\"/p/blog-will-ai-replace-the-artists/images/hugging-face-treading.png\"\u003e\n\u003cimg loading=\"lazy\" decoding=\"async\" src=\"/p/blog-will-ai-replace-the-artists/images/hugging-face-treading_hu14304970296204830138.webp\" alt=\"Popular image generation models on Hugging Face\" width=\"300\" height=\"581\" srcset=\"/p/blog-will-ai-replace-the-artists/images/hugging-face-treading_hu12273577464729043496.webp 354w\" sizes=\"(max-width: 328px) calc(100vw - 28px), 300px\"\u003e\n\u003c/a\u003e\n\u003c/figure\u003e\n\n\u003cp\u003eThe images generated by these diffusion models are equally stunning. Whether landscapes, portraits, or paintings in various styles, the generated images exhibit extremely high quality. Here are a few examples generated by these models.\u003c/p\u003e","title":"Will AI Replace Artists' Jobs?"},{"content":" I originally intended to include this section in another article1, but found that this small part contains too much content and is somewhat less relevant to the main text theme, so I decided to extract it into a separate article, which also served as a research dive into diffusion models. Here, I will simply introduce the principles and history of Diffusion models, along with my own integration of related knowledge.\nWhat is a Diffusion Model? For now, let me take a shortcut; for detailed content, you can refer to Professor Li Mu\u0026rsquo;s video2, or the 3 by Su Shen regarding diffusion models, which covers very deep mathematical principles. I will not elaborate further here.\nBut in essence, it generates a complete image from random noise or an initial value. During training, the process is reversed: starting from complete images and gradually transforming them into random noise. The first \u0026lsquo;D\u0026rsquo; in DDPM stands for Denoising.\nA Brief History of Diffusion Models A Brief History of Generative Models Pushed to the Limit\nA Brief History of the Arms Race\n2020.6.19 Ho et al. published the DDPM paper titled 《Diffusion Probabilistic Models for Image Generation》4. 2022.4.13 OpenAI published the Dall·E 2 paper titled 《Hierarchical Text-Conditional Image Generation with CLIP Latents》5. 2022.4.13 Stable Diffusion published the paper 《High-Resolution Image Synthesis with Latent Diffusion Models》6. 2022.5.30 Google Brain released Imagen7 8. 2022.6.19 Google \u0026amp; NVIDIA presented the Tutorial 《Denoising Diffusion-based Generative Modeling: Foundations and Applications》 at CVPR 20229. 2022.8.30 OpenAI published a blog post about image inpainting using Dall·E10. Trial Runs of Various Models Since I am not very familiar with metrics in the field of image generation, and given that different models offer varying parameters, there are differences in generation performance and iteration counts. I have only conducted some trial runs of these models, so this cannot serve as a strict performance metric; it merely illustrates the demo effects provided by each.\nAll images were generated using the same Prompt: On a black starry background, Pikachu stands with a star stick in his right hand\nThe trial addresses for each model are listed below, and these are all free versions:\nDall·E-mini: https://huggingface.co/spaces/dalle-mini/dalle-mini ERNIE-ViLG: https://huggingface.co/spaces/PaddlePaddle/ERNIE-ViLG Stable Diffusion: https://huggingface.co/spaces/stabilityai/stable-diffusion Dream Studio: https://beta.dreamstudio.ai/dream Generation Results:\nGenerated by Dall·E-mini\nGenerated by ERNIE-ViLG\nGenerated by Stable Diffusion\nGenerated by Stability AI\nIn these samples, Baidu\u0026rsquo;s ERNIE-ViLG on Hugging Face produces high-quality images, but offers no way to adjust its parameters, since Baidu\u0026rsquo;s \u0026lsquo;Infinite Exploration\u0026rsquo; service is not publicly available outside Hugging Face. Stability AI, the company behind Stable Diffusion, also produces fairly clear images, although my parameter choices may have been poor, making its results look slightly worse than ERNIE-ViLG\u0026rsquo;s. Dall·E-mini and Dream Studio produced less satisfactory results, but broadly captured the intended meaning.\nRoughly speaking, besides the pre-trained CLIP model, the parameters that have an impact are the number of iterations and the image size. I don\u0026rsquo;t know about other parameters, nor am I very good at tuning them \u0026ndash;.\nReference another article\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nvideo\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nblog\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://arxiv.org/abs/2006.11239\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://arxiv.org/abs/2204.06125\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://arxiv.org/abs/2112.10752\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://imagen.research.google/\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://arxiv.org/abs/2205.11487\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://cvpr2022-tutorial-diffusion-models.github.io/\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://openai.com/blog/dall-e-introducing-outpainting/\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-diffusion-model-trial/","summary":"\u003cblockquote\u003e\n\u003cp\u003eI originally intended to include this section in another article\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e, but found that this small part contains too much content and is somewhat less relevant to the main text theme, so I decided to extract it into a separate article, which also served as a research dive into diffusion models. Here, I will simply introduce the principles and history of Diffusion models, along with my own integration of related knowledge.\u003c/p\u003e","title":"Trial Run of Diffusion Models"},{"content":"Introduction To state my conclusion upfront: I hold a pessimistic view of the domestic open-source environment.\nThis stems from a post on V2EX: 【Domestic Open-Source Environment】- V2EX1, and I\u0026rsquo;ll also share some of my thoughts on open-source projects.\nWhat is Open Source? Open source originally originated from the free software movement2. Note that the free here does not mean \u0026ldquo;free\u0026rdquo; as in gratis, but rather \u0026ldquo;freedom.\u0026rdquo; Free software refers to the freedom to run, study, modify, and share (distribute copies, whether modified or not) the software.\nOpen source refers to a computer program whose source code is available for public use or modification of its original design. The code is released under software license terms. According to these terms, others can download, modify, and release their versions (forks) back to the community3.\nAlright, now let\u0026rsquo;s take a look at what the domestic \u0026ldquo;largest\u0026rdquo; \u0026ldquo;open-source community\u0026rdquo; Gitee is actually like?\n[How to view the May 18th Gitee repository open-source audit, with some already open-source repositories temporarily closed, to be reopened after passing the audit? - Zhihu] 4\nOf course, I registered with Gitee quite early. Not using it doesn\u0026rsquo;t mean it\u0026rsquo;s not good; after all, in terms of speed, it\u0026rsquo;s definitely stronger domestically than GitHub, but this is largely due to some personal bias of mine. Mainly, there are too few projects, and the quality is too poor. The functionality at the time was average, and there were no features like GitHub Pages back then. It was also always, always many steps slower than GitHub. Another point is the various restrictions on repositories, such as capacity limits. I also know that Gitee officially answered this, saying this move was out of helplessness5. Since it has been prohibited by policy, the so-called \u0026ldquo;free software\u0026rdquo; no longer exists.\nDomestic Open-Source Projects First, let me share my understanding of the term 开源项目. I believe the most fundamental requirements for an open-source project are README and LICENSE. These two files form the foundation of an open-source project, directly explaining how the project can be used and distributed. Meanwhile, Awesome-List should not be considered part of the open-source project scope; at best, it counts as open-source documentation. A complete open-source project should have a relatively complete workflow and standardized contribution methods, such as CONTRIBUTION.md, which tells the community how to contribute and submit code, along with basic coding standards. Most importantly, there should be complete documentation and version numbers; even better would be milestone or roadmap, so that the community can better participate and see the future development plans of the open-source project. Of course, someone must also be available to address issues in issues and Discussion. Additionally, there should be some basic test cases and CI/CD configurations to ensure project quality.\nI have to admit that there are actually many excellent open-source projects in China, with many being incubated by the Apache Foundation.\nBut if you observe, you\u0026rsquo;ll find that these excellent open-source projects are almost entirely undertaken by enterprises, while excellent projects by individual developers are extremely rare.\nOpen source in China belongs to enterprises; there are no individual developers. There is open-source software, but no open-source community.\nEnvironment for Individual Developers In China, an open-source project has two possible fates: one is being acquired by an enterprise (which I believe is the best outcome for an open-source project), and the other is the creator founding their own company.\nLet\u0026rsquo;s discuss the first point first. Of course, most enterprises are unwilling to do this, as it represents a significant expense. What benefits does it bring to the enterprise? Since it\u0026rsquo;s open source, it can already be used freely, so\u0026hellip; the difference between acquiring and not acquiring is merely making the roadmap better aligned with the enterprise\u0026rsquo;s project needs. Moreover, in the current domestic pandemic environment, layoffs and \u0026ldquo;graduations\u0026rdquo; have become commonplace. Funds and personnel previously allocated to internal open-source departments within enterprises have mostly been shifted to business departments, making acquisition even less likely.\nIf the first path is not viable, what about the second? The difficulty is self-evident: how to achieve profitability? Naturally, it would be 2B (business-to-business), integrating with enterprise APIs. This is essentially similar to the first point. As for how to execute 2B, that\u0026rsquo;s where the real challenge lies. What you need to do is use your open-source project as your business card, showcase it to companies, and then proceed with cooperation. However, as an open-source project, if you want to maintain this business card and ensure its long-term survival, you will inevitably need to go 2C (business-to-consumer).\nIs there any other path besides focusing solely on 2C? Yes, Sponsor. Oh, right, I forgot to mention that developers in mainland China cannot register as Sponsors on GitHub. This means you can only achieve sponsorship by placing your own QR code, and you won\u0026rsquo;t be able to use many of GitHub\u0026rsquo;s Sponsor features. Gitee seems to lack a Sponsor system entirely; the domestic review process for this will likely take some time, given it involves tax issues. However, Sponsorships can still provide some help. In reality, whether domestic or international, the outcome is often similar. The truly effective approach should be to use a Pro version to create differentiation and guide users to pay. I believe this is the best path an individual developer can take on their own.\nI wonder how many people still remember the colors.js incident6. The author chose the MIT License, hoping to generate revenue or job opportunities through Sponsorships. However, the result was that enterprises used it, but no one paid. Eventually, the author injected malicious code, causing numerous projects to malfunction, and the author\u0026rsquo;s GitHub account was even temporarily banned. As a developer, I can easily understand the author\u0026rsquo;s feelings, but I don\u0026rsquo;t understand why they chose such an open-source license in the first place. If you want to profit once, either create differentiation or simply avoid using such an unrestricted open-source license. hcaptcha-challenger This project is similar; many people use it to scrape Discord tokens or engage in gray-market activities to make money, yet no one is willing to open-source their work or comply with the GPLv3.0 License provided by the project.\nChina is gradually strengthening its legal framework regarding open-source licenses. There have been legal disputes arising from open-source licenses that have been adjudicated, demonstrating the validity of open-source licenses under Chinese law. I hope that similar incidents triggered by open-source licenses can be properly resolved within China\u0026rsquo;s open-source projects.\nUser Environment A very simple question: if you are not a high-intensity GitHub user (like me, who browses GitHub more than their Qzone and WeChat Moments\u0026hellip;), visiting GitHub is usually purposeful, such as looking for experimental code or finding existing projects to use. It\u0026rsquo;s not common for beginners to use it smoothly (after all, you need to bypass the Great Firewall to download releases and view images in READMEs), let alone expect them to comply with open-source licenses. They might not even know what an open-source license is. Of course, some beginners not only know nothing but also refuse to read the code, directly opening issues to ask questions without checking the Wiki, and they have a terrible attitude, as if open-source developers exist solely to serve them.\nThis is evident from the hcaptcha-challenger project. Myself and another developer frequently receive emails \u0026lsquo;greeting\u0026rsquo; us. Why should I solve problems unrelated to the project for you? For instance, I once received an email asking me to help solve a CAPTCHA for a specific website. Let alone the fact that the project\u0026rsquo;s original intent was to combat hCaptcha, not target their \u0026lsquo;specific website,\u0026rsquo; the code is open-source, yet they are unwilling to even understand the project structure, purely seeking results. I certainly won\u0026rsquo;t help people like this. I don\u0026rsquo;t believe that someone who needs translation software just to reply to an email \u0026lsquo;greeting\u0026rsquo; me can bring any future value. It\u0026rsquo;s inherently a gray industry, and I have no desire to get involved.\nOf course, this isn\u0026rsquo;t specific to the domestic environment, so I might be slightly off-topic, but many domestic users share the same attitude. For example, the controversy surrounding YOLOv7 made me feel embarrassed (I can\u0026rsquo;t find it now, but someone used Chinese to verbally attack the original YOLOv7 project\u0026rsquo;s naming in an abusive tone). Therefore, I advise everyone to use English when communicating on GitHub. Do not reply using web translation plugins. When asking questions, communicate politely and use honorifics. Respect the developers\u0026rsquo; work; developers are not your employees!!!\nConclusion That\u0026rsquo;s basically the entire content of this article. I sincerely hope that China can build a robust open-source community, at the very least ensuring that it doesn\u0026rsquo;t hinder product and project iterations, and that audits do not become a burden on developers, limiting the development of open-source projects. Additionally, I hope to see a well-established Sponsorship system, allowing excellent individual developers to pursue their interests, maintain their open-source communities, and support themselves financially.\nReference 【Domestic Open-Source Environment】- V2EX\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nfree software movement\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://en.wikipedia.org/wiki/Open_source\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.zhihu.com/question/533388365\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.zhihu.com/question/533388365/answer/2491172345\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://zhuanlan.zhihu.com/p/456125379\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-talk-about-the-open-source-environment-in-china/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eTo state my conclusion upfront: I hold a pessimistic view of the domestic open-source environment.\u003c/p\u003e\n\u003cp\u003eThis stems from a post on V2EX: 【Domestic Open-Source Environment】- V2EX\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e, and I\u0026rsquo;ll also share some of my thoughts on open-source projects.\u003c/p\u003e\n\u003ch2 id=\"what-is-open-source\"\u003eWhat is Open Source?\u003c/h2\u003e\n\u003cp\u003eOpen source originally originated from the free software movement\u003csup id=\"fnref:2\"\u003e\u003ca href=\"#fn:2\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e2\u003c/a\u003e\u003c/sup\u003e. Note that the \u003ccode\u003efree\u003c/code\u003e here does not mean \u0026ldquo;free\u0026rdquo; as in gratis, but rather \u0026ldquo;freedom.\u0026rdquo; Free software refers to the freedom to run, study, modify, and share (distribute copies, whether modified or not) the software.\u003c/p\u003e","title":"A Brief Look at the Open-Source Environment in China"},{"content":"Preface Course selection has just ended, and classes are about to begin, but managing the schedule remains a pain point. Since I just switched phones and am not yet familiar with schedule management, I decided to dig deeper and have summarized several solutions here.\nFirst, based on export settings, methods are divided into exporting from the course website and other methods. For imports, this guide primarily explains the iOS approach; theoretically, Android should be even simpler. Finally, I\u0026rsquo;ll mention some all-in-one solutions, i.e., schedule apps.\nThe content may be incomplete; feel free to contribute~\nICS File Format An ICS file is a plain text file containing some core object, such as DTSTAMP, DTSTART, DTEND, etc., and other 1.\nExporting Course Schedules Exporting from the Course Website First, enter the SEP - Course Selection System.\nIn \u0026lsquo;Select Courses - Select Courses - Selected Courses\u0026rsquo;, click \u0026lsquo;Add to Course Website\u0026rsquo;.\nThen enter the SEP - Course Website.\nClick \u0026lsquo;Schedule\u0026rsquo; on the left to view all schedule information.\nClick \u0026lsquo;Publish (Private)\u0026rsquo; as needed to obtain the ICS file and subscription link.\nOr click \u0026lsquo;Publish (Public)\u0026rsquo; to directly download the ICS file, which supports iCal export.\n(If you\u0026rsquo;re interested in my schedule, feel free to try these links 233)\nTampermonkey Plugin (Recommended) Another method is to use a Tampermonkey plugin written by a senior student to export ICS files. When exporting from the course website, courses are split separately, meaning each class session becomes an individual event rather than a single integrated schedule, which may cause multiple reminders.\nThe script written by the senior student can be found at 2. First, you need to install the Tampermonkey plugin in Chrome, Edge, or any browser you use, then click this link 3 to start the installation.\nAfter installation, go to the Course Selection System - View Personal Schedule, i.e., the link 4, where you can see the plugin. For 2022 Yanqi Lake graduate students, configure it as follows.\nWakeup Schedule Export The Wakeup app supports ICS file export, but unfortunately, the iOS version requires payment, while the Android version is free. So, you just need an Android device. See the section below for importing ICS into the iOS Calendar.\nImporting Course Schedules (iOS) This section mainly covers iOS devices, as many Android devices already support ICS file import. However, iOS does not offer this feature well; one could even say it\u0026rsquo;s quite \u0026lsquo;yaxi\u0026rsquo;.\nHere, I\u0026rsquo;ve roughly learned two solutions.\nImport via Email If you\u0026rsquo;ve already used the iOS Mail app and set up your email, you can send the ICS file to your phone via QQ or WeChat, then forward it to yourself via email. Clicking the ICS attachment will import it into your calendar.\nImport via AirDrop This is quite simple: just AirDrop the ICS file to another device to import it.\nHere\u0026rsquo;s a bit of a pain point: you cannot specify which calendar to import into. If importing many at once, it defaults to the first one. When modifying, you can only change the currently selected one, not all of them at once.\nImport via Outlook If you\u0026rsquo;re using a Windows computer, double-click the ICS file or open Outlook to import it. You\u0026rsquo;ll need to download Outlook on your iOS device and log in, then add the Outlook calendar to your system calendar.\nImport via Subscription I haven\u0026rsquo;t found a good tool to convert ICS files into subscription links yet; maybe I\u0026rsquo;ll write one when I have time?\nSee the next section for a better solution.\nSo, compromise with the course website: copy the subscription link from the course website and add it as a subscribed calendar below (I\u0026rsquo;m not sure how the update status of this subscription calendar works; I don\u0026rsquo;t know if the schedule will become invalid if the link expires. If it doesn\u0026rsquo;t, perhaps this conversion only needs to be done once?).\nOf course, you can also use iOS\u0026rsquo;s built-in sharing feature: click the \u0026lsquo;i\u0026rsquo; icon next to the calendar to share a calendar imported from iCloud for others to subscribe to.\nStoring Your Own Schedule via GitHub (Recommended) Through experimentation, I found that since the schedule is in plain text format, it can be easily hosted on GitHub for iterative updates. Meanwhile, GitHub Pages can make your schedule publicly accessible on the internet. This way, whenever your course schedule changes, you only need to update the ICS file on GitHub.\nThe method is simple: first, create a new GitHub repository, clone it, create a docs folder, place the ics file inside that folder, and commit the changes. We assume the ics file here is course-schedule.ics.\nThen, as shown in the image below, click the 5th option to publish your schedule. The address is the URL in the box followed by your filename; for example, it should be webcal://www.bj-yan.top/webcal/course-schedule.ics. (Note: change the protocol header to webcal.)\nFinally, import the subscribed calendar on your device.\nSchedule Import (Windows) Double-click to open the .ics file, log in to your account, and import it into Outlook.\nOther Schedule Import Options Of course, this method is quite simple; apps like Wakeup, Super Course Schedule, etc., can handle it. But I haven\u0026rsquo;t found a corresponding academic affairs interface? Maybe I\u0026rsquo;m doing it wrong, but since there aren\u0026rsquo;t many classes anyway, it\u0026rsquo;s easier to just do it manually.\nThe most recommended option is USTC Online. Anyway, you\u0026rsquo;ll need to download MOOCs and other platforms anyway; the app includes a schedule, but updates may be delayed—probably taking about a day to refresh.\nPersonally, I prefer the system calendar because it\u0026rsquo;s convenient for reminders, and since it\u0026rsquo;s built-in, managing it is much easier.\nReference https://en.wikipedia.org/wiki/ICalendar\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://github.com/LinHeLurking/Sep_Calendar\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://github.com/LinHeLurking/Sep_Calendar/raw/main/sel.user.js\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://jwxk.ucas.ac.cn/course/personSchedule\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-ucas-course-schedule/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eCourse selection has just ended, and classes are about to begin, but managing the schedule remains a pain point. Since I just switched phones and am not yet familiar with schedule management, I decided to dig deeper and have summarized several solutions here.\u003c/p\u003e\n\u003cp\u003eFirst, based on export settings, methods are divided into exporting from the course website and other methods. For imports, this guide primarily explains the iOS approach; theoretically, Android should be even simpler. Finally, I\u0026rsquo;ll mention some all-in-one solutions, i.e., schedule apps.\u003c/p\u003e","title":"UCAS Course Schedule Solution"},{"content":"Disclaimer For technical sharing only. Users should be aware of technical risks and capable of taking responsibility for their own actions. This tutorial and the author assume no liability.\nPreparation First, a note on the preparation phase: most computers now support IPv6. If you open https://ipv6-test.com/ and do not see IPv6, it is likely because it is not enabled. You can try following the steps below to enable it:\nOpen \u0026ldquo;Network and Internet\u0026rdquo; settings Click \u0026ldquo;Change adapter options\u0026rdquo; Select the corresponding network and open \u0026ldquo;Properties\u0026rdquo; 4. Check \u0026ldquo;Internet Protocol Version 6 (TCP/IPv6)\u0026rdquo; and click OK to save.\nUsing the Proxy Refer to 1. The principle of this method is to use an IPv6 proxy to forward our traffic via IPv6 to the proxy, which then accesses the target services via IPv4/IPv6 and returns the response.\nFor the server configuration selected at Vultr, refer to the figure below.\nHere, the cost is $5 per month for 1 TB of traffic, which is elastic and can be destroyed and rebuilt. If several people share the cost, adding $1 allows an upgrade to 2 TB per month, and it seems the memory might also increase slightly. After purchase, you can find the root username and password under Products.\nHowever, note that many IPs at Vultr have been banned due to abuse, so Google Scholar might be inaccessible. However, you can destroy and rebuild the instance without extra cost. After all, it takes about 5 minutes to set up.\nIf possible, I recommend trying the hysteria protocol; it is quite convenient, though I haven\u0026rsquo;t studied it in detail yet.\nAdditionally, proxy rules can be set to bypass the LAN. However, if domestic traffic is also proxied, you may be unable to send or receive messages unless you manually configure the proxy for QQ.\nDNS64 This method is free but has limitations. However, it is sufficient to bypass Bilibili data usage, which saves the bulk of the data consumption hhh.\nThe bypass is mainly achieved via DNS64. Refer to 2 and 3. Simply put, IPv4 server addresses are resolved into IPv6 addresses. When a request for an AAAA record is made, only the A record and an address pointing to a server that performs IPv4-to-IPv6 conversion are returned, allowing access to IPv4 services via IPv6 addresses.\nThe method is simple: just double-click the IPv6 checkbox mentioned above to open it and manually configure the DNS to a DNS64 address.\nHere I copy the addresses from 2:\nProvider Country/City DNS64 Service NAT64 Prefix Kasper Dupont Finland/Helsinki 2a01:4f9:c010:3f02::1 2a01:4f9:c010:3f02:64::/96 Trex Finland/Tampere 2001:67c:2b0::4 2001:67c:2b0:db32::/96 Trex Finland/Tampere 2001:67c:2b0::6 2001:67c:2b0:db32:0:1::/96 level66.network Germany/Frankfurt am Main 2a09:11c0:f1:bbf0::70 2a09:11c0:f1:be00::/96 Kasper Dupont Germany/Nuremberg 2a01:4f8:c2c:123f::1 2a01:4f8:c2c:123f:64::/96 go6Labs Slovenia 2001:67c:27e4:15::6411 2001:67c:27e4:642::/96 go6Labs Slovenia 2001:67c:27e4::64 2001:67c:27e4:64::/96 go6Labs Slovenia 2001:67c:27e4:15::64 2001:67c:27e4:1064::/96 go6Labs Slovenia 2001:67c:27e4::60 2001:67c:27e4:11::/96 Kasper Dupont Netherlands/Amsterdam 2a00:1098:2b::1 2a00:1098:2b::/96 Tuxis Netherlands/Central 2a03:7900:2:0:31:3:104:161 2a03:7900:6446::/96 Kasper Dupont UK/London 2a00:1098:2c::1 2a00:1098:2c::/96 Multi-device Support Sometimes the campus network is unstable, and smart devices like phones may experience disconnections. You can try enabling your computer\u0026rsquo;s mobile hotspot to share the internet with your phone, then configure a proxy IP on the phone to achieve network sharing.\nI encountered an issue where the phone couldn\u0026rsquo;t connect to the hotspot; simply enable sharing in the properties.\nTo set up a WiFi proxy on iOS, tap the \u0026lsquo;i\u0026rsquo; icon on the right, scroll to the bottom where you\u0026rsquo;ll find \u0026lsquo;Proxy\u0026rsquo;, and set the proxy to 电脑ip:端口.\nTo check your computer\u0026rsquo;s IP, use ipconfig /all to find the address corresponding to your mobile hotspot. If you\u0026rsquo;re using macOS or another Linux system, you can use ip addr or ifconfig for the query.\nIf you\u0026rsquo;re using v2RayN, first check 设置-参数设置 to enable 允许来自局域网的连接. Your port should then be 10809. If your v2RayN version is newer, it might have changed to 10811.\nIf you\u0026rsquo;re using SSR, open 右键小飞机-选项设置-本地代理 and enable 允许来自局域网的连接. You should see the port listed as 1080; just fill that in.\nHere\u0026rsquo;s a CDN attached: https://dl.capoo.xyz/\nReference https://github.com/Tremb1e/ucas_ipv6_bypass\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttp://blog.cloudwai.com/archives/201/\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://www.whosneo.com/free-by-ipv6/\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-ucas-ipv6-bypass/","summary":"\u003ch2 id=\"disclaimer\"\u003eDisclaimer\u003c/h2\u003e\n\u003cp\u003eFor technical sharing only. Users should be aware of technical risks and capable of taking responsibility for their own actions. This tutorial and the author assume no liability.\u003c/p\u003e\n\u003ch2 id=\"preparation\"\u003ePreparation\u003c/h2\u003e\n\u003cp\u003eFirst, a note on the preparation phase: most computers now support IPv6. If you open \u003ca href=\"https://ipv6-test.com/\"\u003ehttps://ipv6-test.com/\u003c/a\u003e and do not see IPv6, it is likely because it is not enabled. You can try following the steps below to enable it:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003eOpen \u0026ldquo;Network and Internet\u0026rdquo; settings\u003c/li\u003e\n\u003c/ol\u003e\n\n\u003cfigure class=\"align-center\"\u003e\u003ca href=\"/p/blog-ucas-ipv6-bypass/images/step1.png\"\u003e\n\u003cimg loading=\"lazy\" decoding=\"async\" src=\"/p/blog-ucas-ipv6-bypass/images/step1_hu8493768318500549173.webp\" alt=\"Windows Network and Internet Settings\" width=\"498\" height=\"196\" srcset=\"/p/blog-ucas-ipv6-bypass/images/step1_hu9208066236628573204.webp 360w, /p/blog-ucas-ipv6-bypass/images/step1_hu17331624302947917047.webp 480w, /p/blog-ucas-ipv6-bypass/images/step1_hu8493768318500549173.webp 498w\" sizes=\"(max-width: 526px) calc(100vw - 28px), 498px\"\u003e\n\u003c/a\u003e\n\u003c/figure\u003e\n\n\u003col start=\"2\"\u003e\n\u003cli\u003eClick \u0026ldquo;Change adapter options\u0026rdquo;\u003c/li\u003e\n\u003c/ol\u003e\n\n\u003cfigure class=\"align-center\"\u003e\u003ca href=\"/p/blog-ucas-ipv6-bypass/images/step2.png\"\u003e\n\u003cimg loading=\"lazy\" decoding=\"async\" src=\"/p/blog-ucas-ipv6-bypass/images/step2_hu6346024564828873715.webp\" alt=\"Windows Network Adapter Options\" width=\"720\" height=\"472\" srcset=\"/p/blog-ucas-ipv6-bypass/images/step2_hu13256698088982925233.webp 360w, /p/blog-ucas-ipv6-bypass/images/step2_hu2650914821003701998.webp 480w, /p/blog-ucas-ipv6-bypass/images/step2_hu6346024564828873715.webp 720w, /p/blog-ucas-ipv6-bypass/images/step2_hu8959051285318074124.webp 1080w, /p/blog-ucas-ipv6-bypass/images/step2_hu10153074170135858454.webp 1404w\" sizes=\"(max-width: 748px) calc(100vw - 28px), 720px\"\u003e\n\u003c/a\u003e\n\u003c/figure\u003e\n\n\u003col start=\"3\"\u003e\n\u003cli\u003eSelect the corresponding network and open \u0026ldquo;Properties\u0026rdquo;\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003e\n\u003cfigure class=\"align-center\"\u003e\u003ca href=\"/p/blog-ucas-ipv6-bypass/images/step3.png\"\u003e\n\u003cimg loading=\"lazy\" decoding=\"async\" src=\"/p/blog-ucas-ipv6-bypass/images/step3_hu5863206460685274828.webp\" alt=\"Open Network Adapter Properties\" width=\"569\" height=\"385\" srcset=\"/p/blog-ucas-ipv6-bypass/images/step3_hu17312083354253460997.webp 360w, /p/blog-ucas-ipv6-bypass/images/step3_hu3965365706436947280.webp 480w, /p/blog-ucas-ipv6-bypass/images/step3_hu5863206460685274828.webp 569w\" sizes=\"(max-width: 597px) calc(100vw - 28px), 569px\"\u003e\n\u003c/a\u003e\n\u003c/figure\u003e\n 4. Check \u0026ldquo;Internet Protocol Version 6 (TCP/IPv6)\u0026rdquo; and click OK to save.\u003c/p\u003e","title":"UCAS Bypasses Campus Network Data Billing Using IPv6"},{"content":"I I don\u0026rsquo;t know why lately I seem to dislike using spaces and commas; even with such a long sentence, not adding any punctuation makes me feel like I can\u0026rsquo;t catch my breath.\nII Friends who found jobs are all heading to work; it feels rare to still keep in touch with friends after graduation. The life of meeting up every day at school to play games has turned into a life where we just wait until evening to play a bit and chat.\nI\u0026rsquo;m also a bit confused about my future path; I\u0026rsquo;m not quite sure what I really want. The root cause seems to be issues of time and age.\nIII Sometimes I feel like I\u0026rsquo;m busy all day, but looking back, I don\u0026rsquo;t know what I\u0026rsquo;ve actually done day by day. It\u0026rsquo;s strange.\nFor example, I feel like I haven\u0026rsquo;t improved at all since finishing my recommendation for graduate school, but thinking back, I seem to have learned quite a bit. Where did that learning go?\nI feel there\u0026rsquo;s a misconception in my learning approach; maybe my foundation in subjects like advanced calculus and linear algebra is too weak (?). Yet sometimes when I flip through them, I can understand them, but I just forget them. This shows how important it is to build my own knowledge base, but organizing these notes also takes a lot of time, which is a bit confusing. Has my memory deteriorated?\nIV Regarding family, I don\u0026rsquo;t know how to comment, but I just feel my mom is overly \u0026rsquo;enthusiastic,\u0026rsquo; which has led to many arguments between us.\nI feel a good state is: appearing when needed, offering support, having a harmonious relationship, and not disturbing each other.\nV Sometimes I realize I\u0026rsquo;m very willing to rewrite a project from scratch; this might be the reason I have so many small projects.\nBut when it comes to projects I should be working on, like the update to my FedHF framework, I always keep putting it off.\nThere are probably two reasons: besides the lack of feedback, subjectively I don\u0026rsquo;t really want to do it. For example, regarding that configuration file issue, I actually know how simple it is\u0026hellip; or if it\u0026rsquo;s complex, I\u0026rsquo;m even less willing to change it, and I\u0026rsquo;d rather just rewrite it from scratch.\nAm I just born to like being pushed?\nVI Regarding parting, it\u0026rsquo;s truly only at the very end, when separation is inevitable, that one feels unforgettable pain, like that most lucid 4:30 AM.\nVII After communicating more with friends in the industry, I start to feel the question of meaning and whether something can create social value, and my thinking undergoes a slight shift.\nVIII Sometimes I really want to expand my influence; I will do so in the future, not limited to GitHub, but also including other self-media platforms.\n","permalink":"https://blog.bj-yan.top/en/p/misc-20220706/","summary":"\u003ch2 id=\"i\"\u003eI\u003c/h2\u003e\n\u003cp\u003eI don\u0026rsquo;t know why lately I seem to dislike using spaces and commas; even with such a long sentence, not adding any punctuation makes me feel like I can\u0026rsquo;t catch my breath.\u003c/p\u003e\n\u003ch2 id=\"ii\"\u003eII\u003c/h2\u003e\n\u003cp\u003eFriends who found jobs are all heading to work; it feels rare to still keep in touch with friends after graduation. The life of meeting up every day at school to play games has turned into a life where we just wait until evening to play a bit and chat.\u003c/p\u003e","title":"Recent Reflections - 20220706"},{"content":"Prologue It feels as though four years have passed in the blink of an eye. I once wrote an article 1 to share my academic life in college. Of course, university life wasn\u0026rsquo;t just about academics; it also included the unforgettable daily experiences.\nTrying to recall the entirety of my university life is always complicated. It\u0026rsquo;s hard to remember every single event; even significant moments often result in fragmented, scattered recollections. Perhaps only sitting with close friends, recalling things together, can truly bring joy. So here, I\u0026rsquo;ll just speak freely, saying whatever comes to mind.\nPart One In the twinkling of an eye, four years of university life are coming to a close, and I am about to leave the little island where I\u0026rsquo;ve lived for four years. To be honest, I feel a bit melancholy.\nI remember before coming to Hainan, I was always anxious about its climate. I\u0026rsquo;ve always been someone who fears heat, perhaps because I had enough of stifling classrooms during high school. Back then, in a high school classroom of merely a few dozen square meters, over fifty of us sat tightly packed together. Sometimes, even in summer, the air conditioning wouldn\u0026rsquo;t be turned on, especially as the college entrance exam approached, with the excuse of \u0026ldquo;fearing students might catch a cold.\u0026rdquo; Whenever this happened, especially during evening self-study, stepping out of the classroom to breathe fresh air and feel the cooler outdoor temperature always felt incredibly refreshing. That\u0026rsquo;s why I always thought of myself as someone who \u0026ldquo;fears heat.\u0026rdquo;\nBut when I truly arrived here, aside from the scorching sun during military training, I found that I actually prefer this climate. You might not imagine it, but in this province at the southernmost tip of the motherland, in a tropical region, sometimes, even on sunny days, the temperature in Hainan is significantly lower than on the mainland. And after rain—usually in Hainan, it rains around noon or afternoon—the temperature truly feels comfortable. Moreover, if you follow the locals\u0026rsquo; daily rhythm—morning tea in the morning, a long nap at noon, afternoon tea in the late afternoon, and a late-night snack in the evening—Hainan is truly a very livable place.\nOf course, it cannot be denied that Hainan is indeed very \u0026ldquo;sunny,\u0026rdquo; but with proper sun protection, it\u0026rsquo;s actually quite manageable. Even guys need to carry umbrellas; otherwise, after just a few minutes of exposure, the neck can feel a burning, painful heat.\nAnother point that must be mentioned is Hainan\u0026rsquo;s sky; I truly love it. The clouds in Hainan are incredibly low. In some small mountainous areas, like Wuzhishan, you can see a sea of clouds, and you\u0026rsquo;re actually above the cloud layer. Keep in mind that the main peak of Wuzhishan is only 1,867 meters above sea level. There are also the colorful auspicious clouds at dusk, some like flames, or double rainbows after a storm\u0026hellip;\nThinking about it, retiring in Hainan in the future might actually be a pretty good idea (XD).\nPart Two This is about graduation.\nUnlike the hurried graduations of high school and middle school—where everyone took a \u0026ldquo;life-defining exam\u0026rdquo; and then went their separate ways; I didn\u0026rsquo;t even feel much when someone else picked up my documents upon returning to campus—graduation from university was a process to be enjoyed.\nStarting from the second semester of junior year, the course load began to decrease. By the first semester of senior year, everyone was busy with their own pursuits: internships, postgraduate entrance exams, or recommended admissions. By the second semester of senior year, apart from the final thesis, everyone was striving for their respective futures. When the thesis was completed and the final file bag submitted, university life was nearing its end. Stepping out of the school gate for the last time marked the true graduation.\nOne interesting aspect of enjoying graduation is that everything you do now is done one less time; in other words, every instance is a specific \u0026ldquo;last time.\u0026rdquo; The last time visiting a certain cafeteria or a specific food stall, the last time walking down this path on campus, the last time crossing the century-old bridge, the last time\u0026hellip; the last time eating and chatting with friends from different groups, observing how everyone has changed over the four years.\nPart Three This is about Hainan University.\nIn previous articles, I\u0026rsquo;ve mentioned how my perception of Hainan University has changed. I believe the reason is that our cohort was fortunate to arrive at a pivotal time; the university\u0026rsquo;s tremendous changes and development were unfolding right before our eyes.\nAir conditioning in Hainan University dorms actually started with our cohort; we must thank the efforts of the seniors and teachers before us. It is said that before air conditioning was installed, seniors and juniors had to take multiple showers a day, and at night, some even slept in the Siyuan Lecture Hall. Fortunately, I never had to experience that (x).\nAdditionally, the arrival of President Luo Qingming, with his title of Academician and the resources he brings, has brought about significant improvements to Hainan University. Taking our school as an example, another Academician and many professors have joined us, significantly strengthening our faculty, especially in the School of Biomedical Engineering. Moreover, it seems the administrative efficiency of Hainan University has also become extremely high.\nHainan University\u0026rsquo;s ranking is rising rapidly. Undoubtedly, the university is developing in a positive direction. Welcome to apply!! (The value of my diploma still needs the efforts of the juniors!)\nEpilogue Friends, there\u0026rsquo;s no need to say too much. I hope we can share more stories when we meet again next time!\narticle \u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/misc-good-bye-hainan/","summary":"\u003ch2 id=\"prologue\"\u003ePrologue\u003c/h2\u003e\n\u003cp\u003eIt feels as though four years have passed in the blink of an eye. I once wrote an article \u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e to share my academic life in college. Of course, university life wasn\u0026rsquo;t just about academics; it also included the unforgettable daily experiences.\u003c/p\u003e\n\u003cp\u003eTrying to recall the entirety of my university life is always complicated. It\u0026rsquo;s hard to remember every single event; even significant moments often result in fragmented, scattered recollections. Perhaps only sitting with close friends, recalling things together, can truly bring joy. So here, I\u0026rsquo;ll just speak freely, saying whatever comes to mind.\u003c/p\u003e","title":"Farewell, My Little Island"},{"content":"I\u0026rsquo;ve used this technique before in a project; to put it simply, it\u0026rsquo;s quite straightforward: perform data augmentation during testing to enhance the overall system\u0026rsquo;s robustness.\nAdditionally, hCAPTCHA\u0026rsquo;s recent update involves adding noise to images, which is essentially a relatively simple model attack. However, what piques my curiosity is: what is the actual utility of performing a model attack in this specific scenario?\nI believe most model attacks and defenses occur in real-world scenarios, since it\u0026rsquo;s difficult to specifically alter real-world objects. But in this scenario\u0026hellip; I simply don\u0026rsquo;t get it.\nI just need to apply a perturbation to the noise you added, or perform a simple denoising operation—just one line of img = cv2.fastNlMeansDenoisingColored(img, None, 10, 10, 7, 21). At the very least, if I add a filter img = cv2.GaussianBlur(img, (5, 5), 0), it seems your noise-adding operation becomes useless.\nAfter all, my model\u0026rsquo;s input isn\u0026rsquo;t determined by you, so your attack on my model is nearly 0% effective.\nSo, I truly don\u0026rsquo;t understand this noise-adding operation. Is it meant to thwart humans? To make humans\u0026rsquo; eyes blur while robots make correct judgments, so that only robots pass and humans fail? That\u0026rsquo;s just too hilarious 8, haha.\n@hCAPTCHA, please come up with more interesting challenges; lately, everything feels boring. Bring on something difficult!\nI believe the ultimate form of CAPTCHAs will not be image-based ones, but rather some form of environmental monitoring that makes the CAPTCHA invisible to users. This might be the true future form of CAPTCHAs, just as the ultimate goal of distributed systems is to be invisible.\n","permalink":"https://blog.bj-yan.top/en/p/blog-hcaptcha-test-time-augmentation-and-model-attack-defense/","summary":"\u003cp\u003eI\u0026rsquo;ve used this technique before in a project; to put it simply, it\u0026rsquo;s quite straightforward: perform data augmentation during testing to enhance the overall system\u0026rsquo;s robustness.\u003c/p\u003e\n\u003cp\u003eAdditionally, hCAPTCHA\u0026rsquo;s recent update involves adding noise to images, which is essentially a relatively simple model attack. However, what piques my curiosity is: what is the actual utility of performing a model attack in this specific scenario?\u003c/p\u003e\n\u003cp\u003eI believe most model attacks and defenses occur in real-world scenarios, since it\u0026rsquo;s difficult to specifically alter real-world objects. But in this scenario\u0026hellip; I simply don\u0026rsquo;t get it.\u003c/p\u003e","title":"Test-Time Augmentation and Model Attacks/Defenses: Starting from hCaptcha Image Noise"},{"content":"🍕 Easy Configuration(ezkfg) GitHub: 🍕 Easy Configuration(ezkfg)\nI want to talk about a small project I recently wrote, ezkfg.\nWhy this name? Mainly because names like (kfg, config, cfg, ezcfg, ezconfig) were already taken\nSecondly, configuration files are always in high demand; I use them in many projects but never found one I truly liked. Rather than keep searching, I decided to build my own wheel, taking the best parts and discarding the rest.\nCome quickly to pip install ezkfg!\nSeveral Configuration Methods argparse:\nThe most common is the argparse approach, which supports many data types, allows setting default values, supports setting optional values, and even supports descriptions for optional values. This is also the method I frequently use in many projects, especially in deep learning projects, because it allows for very convenient specification of experimental parameters. However, parameters are hard to persist; you either have to manually write file reading code or pass a list in the file first for parameter parsing.\nconfig.py:\nThis is probably my favorite and most commonly used form besides the above, because it is native Python content. You can define multiple config classes and import them via from config import config_xxx as config, which is quite convenient for configuration while supporting all Python data types. However, since it requires modifying import package information, changing parameters might require editing at least two files each time, which can be confusing.\nYAML:\nYAML is also a very flexible configuration file type; it has a simple structure, can be loaded from files, is very convenient for reproducibility, and supports data structures like lists. It is a configuration file type I really like.\njson:\nSame as above, but the file structure is more complex, and it becomes less intuitive once the content grows.\nini:\nConvenient for selecting multiple configurations, with built-in library support, but it lacks data types; everything is read as strings, requiring manual data conversion. Additionally, it is in k-v form, so it has no hierarchy.\nAt the same time, I also came across the addict project, where nested . attribute access is really quite convenient, but it lacks file I/O.\nClarifying Requirements So, we still need to clarify the requirements.\nThe file is just a carrier; I personally believe that differences in file formats are all acceptable, perhaps with slight variations depending on the application scenario. For example, if I want to send my experimental content to others or quickly reproduce it on another machine, using a configuration file is clearly far superior to other methods.\nAdditionally, the form of invocation in experiments is also very important. First is pre-loading parameters. During pre-loading, some parameters need default values, while others might obviously not be used; in such cases, it is best not to occupy storage and to simplify the weight of the configuration file.\nSecond, data types are necessary, since parameters in experiments can take various forms.\nThen there is parameter reference in the code. For specific parameters, if I want to achieve hierarchical access using nested . instead of calling with a long string of []. However, in other scenarios, such as two methods that share the same hyperparameters, using . for invocation is not conducive to code reuse. A better approach would be the config['method'].alpha form, which looks more intuitive; this way we can dynamically modify method to use the corresponding alpha values. Therefore, both forms are necessary.\nFinally, there is file I/O, which must be supported. If you need to deploy elsewhere, obviously using a configuration file directly without modifying any original project code yields much better results.\nImplementation First and foremost are attribute access and index access; both are mandatory and must support nesting. Index access via [] is nothing special; it is a built-in feature because every class in Python actually has a __dict__ variable that stores all attributes of the class. All configuration parameters also borrow this dictionary for storage.\nFirst is initialization. Currently, it supports several initialization methods: dictionaries, lists, and Config classes. dict requires a recursive solution, while Config can simply use the built-in update method.\nAmong these, __getitem__ and __setitem__ methods are implemented by calling __getattr__ and __setattr__ respectively, which is straightforward without ambiguity. The __setitem__ method requires splitting first, then recursing. However, the implementation of __getattr__ is also simple, just recursive querying. But __setattr__ is different; before setting, it calls the get method to query first, and if it does not exist\nFinally, solve the file I/O problem.\nRegarding handler, a registry is provided for users who like customization, allowing them to directly register custom handler into the package, or to parse custom file formats using existing handler.\nRoadmap The project has currently reached the v0.0.1 version, having gone through 3 pre-release versions.\nHowever, there are some issues that are visibly apparent.\nFirst, type conversion is not yet supported; default parameters can only be achieved by pre-loading a default file or by pre-converting argparse. The interfaces for py and ini files still need optimization, as the names are currently hard-coded.\nTakeaways To be honest, this project has indeed taught me a lot about various configuration file formats and their characteristics, as well as Python\u0026rsquo;s built-in functions and their call order. For example, __setattr__ calls __getattr__ first, but this is not implemented within __setattr__; instead, it is implemented at a higher level, so proper handling is required.\nBesides that, the more important thing is to make it convenient for me to write configuration files for my own projects, since it\u0026rsquo;s likely that no one else will use it, hhh\n","permalink":"https://blog.bj-yan.top/en/p/blog-project-ezkfg/","summary":"\u003ch2 id=\"-easy-configurationezkfghttpsgithubcomyyyanbjezkfg\"\u003e\u003ca href=\"https://github.com/yyyanbj/ezkfg\"\u003e🍕 Easy Configuration(ezkfg)\u003c/a\u003e\u003c/h2\u003e\n\u003cp\u003eGitHub: \u003ca href=\"https://github.com/yyyanbj/ezkfg\"\u003e🍕 Easy Configuration(ezkfg)\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eI want to talk about a small project I recently wrote, ezkfg.\u003c/p\u003e\n\u003cp\u003eWhy this name? \u003cdel\u003eMainly because names like (kfg, config, cfg, ezcfg, ezconfig) were already taken\u003c/del\u003e\u003c/p\u003e\n\u003cp\u003eSecondly, configuration files are always in high demand; I use them in many projects but never found one I truly liked. Rather than keep searching, I decided to build my own wheel, taking the best parts and discarding the rest.\u003c/p\u003e","title":"ezkfg: Design and Implementation of a Python Configuration Tool"},{"content":"Preface The reason is that I recently have quite a lot of tasks and needed a schedule management software again, but it seems I\u0026rsquo;ve finally found one that suits me well.\nSchedule Management Software That Fails to Meet Needs Looking back to high school, I still kept the habit of writing down homework in a notebook that I started in elementary school (though I slacked off towards the end), but indeed, \u0026ldquo;a good memory is not as good as a bad pen.\u0026rdquo; Writing things down ensures you can find them when you look for them, and besides homework, there was nothing else recorded.\nSince starting university and getting my own phone, I naturally stopped using small notebooks for homework, so I needed a schedule management app on my phone. At first, I just used the phone\u0026rsquo;s built-in Notes app to simply record when the deadlines for various course assignments were. It was okay, but sometimes, without reminders, I would forget to check even if I had written it down, and I wouldn\u0026rsquo;t clear them in time.\nAfter joining some organizations, schedule management became complicated: when to post articles, when there were events, when there were meetings, weekly meetings\u0026hellip; a mess. At this point, Notes was a bit inadequate. After evaluating many software options, I chose TickTick. It had a to-do list, weekly and monthly views, but recurring tasks were a paid feature\u0026hellip; After using it for a while, I quickly gave up. I was too lazy to add recurring tasks like weekly meetings one by one, and accidentally clicking a task meant it took half a day to find it again in the trash bin\u0026hellip; and so on.\nLater, after leaving the Youth Volunteer Association and joining some project teams, the need on my phone actually decreased because most of the time I was working indoors. In this case, it would be best to operate on a computer. After trying some software, I simply used the built-in Sticky Notes on Windows, which essentially functioned as a to-do list. However, it had no time settings, no various views, no recurring tasks; it was just a sticky note. Besides being lazy about cleaning it up, I also stopped updating it frequently.\nLater, I used Google Calendar. To be honest, I found it quite comfortable to use. It integrates with many other apps; for example, if you receive a meeting invitation in your email, you can add it directly to Google Calendar, and in Slack, you can automatically adjust your status based on Google Calendar. It also supports external links, drag-and-drop adjustments, and so on. However, for well-known reasons, it is inconvenient to use once you leave the indoor environment.\nWhat Are My Needs? From small notebooks, Notes, calendars to to-do lists, TickTick, Windows Sticky Notes, Pomodoro timers, Wolai, and Google Calendar, it seems that every schedule management software has failed to fully satisfy me. I started wondering: is this my problem? What exactly are my needs? What are the absolute requirements for my schedule management?\nBefore this, my understanding of my needs was actually quite chaotic. I didn\u0026rsquo;t have a very clear perception of them, always attributing the issues to the software\u0026rsquo;s shortcomings rather than the software failing to meet specific needs of mine.\nHere, I have summarized two parts. Actually, the needs are simple, but they require not just a schedule management software, but a combination of a to-do list and a calendar view for schedule management.\nBecause many tasks do not have specific deadlines, putting them into a calendar view for schedule management can easily cause confusion. Or, placing them only in the calendar might feel unclear, whereas a simple list in Notes is more direct. Schedule management, on the other hand, refers to meetings (things that must be done at a specific date and time, not things that are due at a specific date and time).\nTo-Do List: First and foremost, some tasks have no clear deadline; they are just things I want to do. Some are due on a specific day, and some are due at a specific time on a specific day. They need a list for clear display, making it easy to plan schedules for the coming days. Schedule Management: For this, I prefer having weekly and monthly views. An annual view doesn\u0026rsquo;t seem essential to me, but it might be very useful during annual summaries. Having a heatmap like GitHub\u0026rsquo;s would also be nice? This requires specific meeting times, and ideally, it should support box selection, allowing for quick division and planning of time blocks. At the same time, recurring tasks are a must-have feature. Other: The most important item, and the reason why I am dissatisfied with many mobile schedule management apps, is cross-device synchronization. This is really important; you never want to enter the same content twice. Ideally, it should support push notifications across multiple devices. Another point is the quick creation of tasks or schedules. You never want to spend two minutes tweaking a task here and there to create two separate tasks. Bonus Features: It would be best if I could place it directly on my Blog or other web pages. This serves two purposes: it makes it convenient for colleagues or others to understand my schedule and arrange their work accordingly, and it allows me to edit online. In this way, using the web page achieves cross-device synchronization, and you don\u0026rsquo;t even need to download any client. Feishu — A Temporary Alternative First, I won\u0026rsquo;t casually use small, obscure software. Such apps are often irresponsible with your data and may not exist for the long term, so some of your records might disappear forever if they go out of business.\nFeishu is a product of ByteDance. Being from a major tech company, data security is naturally guaranteed. Another reason I used it is that other projects required Feishu, so I casually looked into it. Surprisingly, I found that it could basically meet my needs: it has tasks and a calendar, supports cross-device synchronization, offers multiple views, and allows box selection\u0026hellip;\nAnother surprising feature is that it actually has an auxiliary time zone function. When I wasn\u0026rsquo;t involved in any cross-border projects, I thought this feature was absolutely useless. But when I actually had the need, I found it truly amazing; it\u0026rsquo;s much better than searching for time differences on Baidu every time.\nThe only regret might be the inability to link tasks with the calendar, so that some tasks can be automatically added to the calendar.\nOpen-Source Software and Serverless Given this need and my heavy daily surfing on GitHub, I naturally thought of using some open-source projects. I have to admit, there are indeed some well-made ones with web views and decent aesthetics. However, they don\u0026rsquo;t provide deployment instructions and require you to purchase their Pro features. Sigh\u0026hellip; That\u0026rsquo;s also understandable; maintaining this data naturally incurs some costs. And if a schedule management app starts showing ads, guess who would still use it?\nServerless itself cannot use databases for storage, and any update creates a new environment, making it impossible to persist data. However, deploying via Serverless is indeed feasible. You don\u0026rsquo;t need to buy a server; you only need an OSS or a database, which might cost just a few dozen yuan per year? You can get a huge amount of storage, which is more than enough for small data like schedules.\nBut the only real challenge is authentication. I haven\u0026rsquo;t figured out how to handle authentication for Serverless yet. Can\u0026rsquo;t we just use GET + Token? Although that would technically work, it looks a bit silly\u0026hellip;\nMaybe in the future, once my frontend coding skills are deep enough, I could build and maintain an open-source project myself?\n","permalink":"https://blog.bj-yan.top/en/p/blog-some-thoughts-on-schedule-management-software/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eThe reason is that I recently have quite a lot of tasks and needed a schedule management software again, but it seems I\u0026rsquo;ve finally found one that suits me well.\u003c/p\u003e\n\u003ch2 id=\"schedule-management-software-that-fails-to-meet-needs\"\u003eSchedule Management Software That Fails to Meet Needs\u003c/h2\u003e\n\u003cp\u003eLooking back to high school, I still kept the habit of writing down homework in a notebook that I started in elementary school (though I slacked off towards the end), but indeed, \u0026ldquo;a good memory is not as good as a bad pen.\u0026rdquo; Writing things down ensures you can find them when you look for them, and besides homework, there was nothing else recorded.\u003c/p\u003e","title":"Choosing Schedule Management Software and Reflections on Personal Workflows"},{"content":"Preface I\u0026rsquo;ve already shared two creations before: Taichi Voxel Challenge 20221\nRecently, after finishing a few urgent tasks at hand, I got the itch to create something new hhh. I had several ideas, but ultimately decided to go with the PVZ theme first.\nThe PVZ project is still a WIP; I plan to refine it further later. If time permits, I might create a few more, but I\u0026rsquo;m about to defend my thesis 555.\nDesign Concept Currently, the PVZ project only features the simplest Peashooter, but I encountered many issues during the process. First, let\u0026rsquo;s analyze the Peashooter\u0026rsquo;s structure!\nThe Peashooter roughly consists of the following parts: the main cannon barrel, eyes, stem, the bud at the back, and the leaves at the bottom.\nMain Cannon Barrel: We can further break this down. It can be viewed as a cylinder, but with curved walls. Seems simple, right? We could directly represent the geometry using SDFs and build it! But I don\u0026rsquo;t know SDFs (x). Actually, this approach has some issues because the muzzle should be positioned lower, so I simply decomposed it into a sphere plus a few circles.\nEyes: No difficulty here. Simply subtract a rectangle from a sphere, place a black rectangle for the eye, and add a white rectangle for the highlight.\n1 2 3 4 def create_eye(p): create_box(p, 2, 8, 3, 0, vec3(0)) create_box(p, 2, 1, 3, 1, vec3(0)) create_box(p + vec3(0, 1, 0), 1, 1, 1, 1, vec3(1)) Stem: This stem\u0026rsquo;s curve stumped me; I don\u0026rsquo;t know Bezier curves. So, I decided to simplify things and just use a sine function hhh, inspired by a work from a certain dragon 233. I defined the amplitude and only selected the parameter range of [0, PI].\n1 2 3 4 5 @ti.func def create_sine_curve(p, A, l, mat, color, dir1=vec3(0, 1, 0), dir2=vec3(0, 0, 1), tk=1): for x, tx in ti.ndrange((0, l + 1), (0, tk)): y = ti.cast(A * ti.sin(1.0 * x / l * ti.math.pi), ti.int32) scene.set_voxel(p + y * dir2 + tx * dir2 + x * dir1, mat, color) The Bud at the Back: Simply draw this as a curve.\nLeaves at the Bottom: The original Peashooter has three leaves at the bottom—two large and one small—but that\u0026rsquo;s too complex. Here, I simplified it by drawing four leaves in four directions. However, leaves are irregular shapes, which are hard to represent\u0026hellip; Then I had a flash of inspiration: draw one leaf ()! This looks very much like the overlapping region of two diagonal semicircles. So, I defined a starting point and direction, then used top-left/bottom-right or top-right/bottom-left as the centers and checked for the overlapping region. But this results in a flat shape. No problem! For the z-coordinate, I used the old method again: simply add two sine functions regarding x and y! (So clever, me)\n1 2 3 4 5 6 7 8 9 10 11 def create_leaf(p, r, dir, mat, color): if dir == 1: for x, y in ti.ndrange((0, r + 1), (0, r + 1)): if x * x + y * y \u0026lt;= r * r and (r - x) * (r - x) + (r - y) * (r - y) \u0026lt;= r * r: z = ti.cast(ti.floor(1 * ti.sin(ti.math.pi * x / r) + ti.sin(ti.math.pi * y / r)), ti.int32) scene.set_voxel(p + vec3(x, z, y), mat, color) elif dir == 2: for x, y in ti.ndrange((0, r + 1), (0, r + 1)): if x * x + (r - y) * (r - y) \u0026lt;= r * r and (r - x) * (r - x) + y * y \u0026lt;= r * r: z = ti.cast(ti.floor(1 * ti.sin(ti.math.pi * x / r) + ti.sin(ti.math.pi * y / r)), ti.int32) scene.set_voxel(p + vec3(x, z, y), mat, color) That concludes the Peashooter. The code to build the Peashooter is as follows:\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 @ti.func def create_peashooter(p): create_ball(p + vec3(0, 18, 0), 8, 1, col_g) for i in range(6): create_circle(p + vec3(0, 16, 6 + i), 3.0, 1, col_g, 1) create_circle(p + vec3(0, 16, 12), 4.0, 1, col_g, 1);create_circle(p + vec3(0, 16, 13), 4.0, 1, col_g, 1) for i in range(8): create_circle(p + vec3(0, 16, 6 + i), ti.max(2.0, ti.min(i - 3.0, 3.0)), 0, col_g, 1) for x, y in ti.ndrange((-1, 1 + 1), (-1, 1 + 1)): create_sine_curve(p + vec3(x, 5, y), 2, 5, 1, col_gd, vec3(0, 1, 0), vec3(0, 0, -1)) create_sine_curve(p + vec3(x, 0, y), 2, 5, 1, col_gd, vec3(0, 1, 0), vec3(0, 0, 1)) create_sine_curve(p + vec3(0, 24, -5), 2, 5, 1, col_gd, vec3(0, 0, -1), vec3(0, 1, 0), 2) create_eye(p + vec3(-3, 20, 5));create_eye(p + vec3(2, 20, 5)) create_leaf(p + vec3(0, 0, -6), 6, 1, 1, col_gdd);create_leaf(p + vec3(-7, 0, 1), 6, 1, 1, col_gdd) create_leaf(p + vec3(-6, 0, -6), 6, 2, 1, col_gdd);create_leaf(p + vec3(1, 0, 1), 6, 2, 1, col_gdd) The rest is the lawn. Due to the grid limit, I couldn\u0026rsquo;t implement the original 6 * 9 setup, nor could I draw the fences on all four sides, so I just created a 6 * 6 scene. After finishing, it felt a bit off, so I added noise around the lawn to fix it.\n1 2 3 4 5 6 7 8 9 @ti.func def create_grass(p, sx, sy, sz, mat, color): create_box(p, sx, sy, sz, mat, color) for x, y in ti.ndrange((0, sx + 1), (0, sy + 1)): if ti.random() \u0026gt; 0.8: scene.set_voxel(p + vec3(0, 0, (ti.random() - 0.5) * 4), mat, color) scene.set_voxel(p + vec3(x, 0, (ti.random() - 0.5) * 4), mat, color) scene.set_voxel(p + vec3((ti.random() - 0.5) * 4, 0, y), mat, color) scene.set_voxel(p + vec3((ti.random() - 0.5) * 4, 0, 0), mat, color) Finally, the line count was slightly over, but after compressing it a bit, it\u0026rsquo;s down to 99 lines!\nConclusion There are a few issues I still haven\u0026rsquo;t resolved \u0026ndash;. First, I don\u0026rsquo;t know how to swap the dimensions of vec3 within @ti.func, for example, turning vec3(x, y, z) into vec3(z, y, x). Otherwise, I could have saved a few lines when building circles, as I was aiming for multi-directional circles, but I ultimately chose the most brute-force if algorithm.\nAnother small detail is that for circles and spheres, to make the circles rounder, I had to relax the boundary conditions, such as x * x + y * y \u0026lt;= r * r + eps. When dealing with spheres, sometimes an extra point appears at the top, so tightening the boundary conditions works better in that case.\nTaichi Voxel Challenge 2022\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-taichi-voxel-challenge-2022-pvz/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eI\u0026rsquo;ve already shared two creations before: \u003ccode\u003eTaichi Voxel Challenge 2022\u003c/code\u003e\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eRecently, after finishing a few urgent tasks at hand, I got the itch to create something new hhh. I had several ideas, but ultimately decided to go with the PVZ theme first.\u003c/p\u003e\n\u003cp\u003eThe PVZ project is still a WIP; I plan to refine it further later. If time permits, I might create a few more, but I\u0026rsquo;m about to defend my thesis 555.\u003c/p\u003e","title":"Taichi Voxel Challenge 2022 - PVZ"},{"content":"As mentioned last time, 1 seaplane, here it comes!\nTo be honest, there\u0026rsquo;s really not much to say, except that there are more weird creatures2\u0026hellip; because there are too many hybrids that don\u0026rsquo;t quite fit any category. I think for the model to understand that it\u0026rsquo;s generating a seaplane, there should be a line underneath. Anyway, I feel the generative model for this task isn\u0026rsquo;t very good; the generated results are too poor.\nAll the code for training and testing is in this commit3. I didn\u0026rsquo;t do much generalization testing, but the results should be decent, given how much data there is.\nInitially, there were many mislabeled samples in the data, which caused significant trouble for the model, leading to overfitting even when the accuracy was relatively high. Alchemy, as it were, does require some experience to find the right stopping point. Perhaps later I can write a Grid Search?\nThe approach this time is actually quite similar to last time, mainly because I directly built a model factory this time. This way, for any new task, I can collect some data, annotate it, train, test, and deploy—all in one go.\nActually, the solution structures for hcaptcha challengers are all quite similar, with almost no changes, except for the elephant one last time which required adding a filter. I\u0026rsquo;ll likely continue building my model factory, with little change unless the model\u0026rsquo;s generalization ability can\u0026rsquo;t keep up, such as with overly detailed images, in which case I might consider modifying the model structure.\nThat\u0026rsquo;s about it. Below is the demo made by @QIN2DIM\nhCaptcha Seaplane Recognition Demo (original GIF) View original GIF As mentioned last time, \u0026#160;\u0026#x21a9;\u0026#xfe0e;\nweird creatures\u0026#160;\u0026#x21a9;\u0026#xfe0e;\ncommit\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-hcaptcha-seaplane/","summary":"\u003cp\u003eAs mentioned last time, \u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e seaplane, here it comes!\u003c/p\u003e\n\u003cp\u003eTo be honest, there\u0026rsquo;s really not much to say, except that there are more weird creatures\u003csup id=\"fnref:2\"\u003e\u003ca href=\"#fn:2\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e2\u003c/a\u003e\u003c/sup\u003e\u0026hellip; because there are too many hybrids that don\u0026rsquo;t quite fit any category. I think for the model to understand that it\u0026rsquo;s generating a seaplane, there should be a line underneath. Anyway, I feel the generative model for this task isn\u0026rsquo;t very good; the generated results are too poor.\u003c/p\u003e","title":"hCaptcha Seaplane Recognition: Model Training and Testing"},{"content":"Voxel Pac-man I came across Taichi\u0026rsquo;s voxel-challenge repository quite a while ago. Since I\u0026rsquo;ve always been a huge fan of pixel art, I naturally kept an eye on voxel art as well; for instance, I previously conducted a deep learning for pixel art1 survey on the topic. While casually browsing the repository, I stumbled upon an issue #1, which got me super excited to try it out myself. The code is available here, and it roughly looks like this:\nOf course, I later found out this was actually an internal beta (x), which explained why I couldn\u0026rsquo;t find any push notifications about this challenge on the official WeChat account for ages.\nOnce the official competition started, I tweaked the content, fine-tuned a few parameters to improve it, added some feed balls, and submitted this Voxel Pac-man. It now has a slightly more sophisticated look (x:\nThe code is about 96 lines long, and the overall logic is quite simple. First, let\u0026rsquo;s look at what pac-man consists of: there\u0026rsquo;s the entire spherical body, but with the mouth open, making it incomplete. The upper and lower parts of the mouth should be considered separate surfaces. Then there are the eyes, and finally, the feed balls.\nI added a lot of customizable parameters at the beginning of the code for easy adjustment. n defines the size of the entire space. Since I hadn\u0026rsquo;t looked at the code myself at the time, I didn\u0026rsquo;t know the space limit was (-64,64), so I just made up a number. The rest involved setting the radius, center point, facing direction, and so on.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 @ti.kernel def initialize_voxels(): n = 60 r = 20 p_center = vec3(0, n // 2, -n // 2) vec_face_to = vec3(0, 0, 1) vec_z = vec3(0, 1, 0) vec_normal = normalize(cross(vec_face_to, vec_z)) mouse_angle = pi / 5 mouse_angle_min = mouse_angle mouse_angle_max = mouse_angle mouse_angle_cos_min = ti.cos(mouse_angle_min) - 0.05 mouse_angle_cos_max = ti.cos(mouse_angle_max) + 0.05 skin_thickness = 0.5 eye_angle = mouse_angle + pi / 18 he_angle = pi / 5 vec_eye_left = rotate(rotate(vec_face_to, vec_normal, -eye_angle), vec_z, he_angle).normalized() vec_eye_right = rotate(rotate(vec_face_to, vec_normal, -eye_angle), vec_z, -he_angle).normalized() p_eye_left = p_center + vec_eye_left * r p_eye_right = p_center + vec_eye_right * r eye_size = 4 The core logic is as follows: first, draw the skin surface by iterating through x, y, and z. If a point is inside the skin, we need to check if it falls within the mouth area. If it does, we skip drawing it; otherwise, we fill it with the skin color. How do we determine if a point is in the mouth? Looking from the side, if the projection of the point onto the central plane forms an angle with the direction of pac-man that falls within our mouth-opening angle range, then it\u0026rsquo;s part of the mouth. Finally, we draw the inside of the mouth. The logic is similar: first, ensure the point is inside the skin, then check if it\u0026rsquo;s within the mouth\u0026rsquo;s central area. If not, we draw it within a certain range above and below the open mouth. Since voxels might not be perfectly standard, we add some parameters to make it look more pleasing.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 for i, j, k in ti.ndrange((-n, n), (-n, n), (-n, n)): x = ivec3(i, j, k) color = vec3(2, 2, 2) # surface if distance(x, p_center) \u0026lt; r + skin_thickness and distance(x, p_center) \u0026gt; r - skin_thickness: # mouse # project to the plane vec_mouse = vec3(i, j, k) - p_center vec_mouse_projected = vec_mouse - vec_normal * dot(vec_mouse, vec_normal) print(vec_mouse_projected) # angle to face to angle_cos = dot(vec_mouse_projected, vec_face_to) / (vec_mouse_projected.norm() * vec_face_to.norm()) if angle_cos \u0026lt;= mouse_angle_cos_max: color = vec3(1, 1, 0.) if distance(vec3(i, j, k), p_eye_left) \u0026lt;= eye_size or distance( vec3(i, j, k), p_eye_right) \u0026lt;= eye_size: color = vec3(0.01, 0.01, 0.01) elif distance(x, p_center) \u0026lt;= r - skin_thickness: # mouse # project to the plane vec_mouse = vec3(i, j, k) - p_center vec_mouse_projected = vec_mouse - vec_normal * dot(vec_mouse, vec_normal) angle_cos = dot(vec_mouse_projected, vec_face_to) / (vec_mouse_projected.norm() * vec_face_to.norm()) if mouse_angle_cos_min \u0026lt;= angle_cos and angle_cos \u0026lt;= mouse_angle_cos_max or vec_mouse_projected.norm( ) \u0026lt; 3: color = vec3(0.0, 0.0, 0.0) if any(color != vec3(2, 2, 2)): scene.set_voxel(vec3(i, j, k), 2, color) Drawing the feed balls is relatively straightforward.\n1 2 3 4 5 6 @ti.func def create_feed_ball(feed_r, feed_p, feed_color): for i, j, k in ti.ndrange((-feed_r, feed_r), (-feed_r, feed_r), (-feed_r, feed_r)): x = ivec3(i, j, k) if distance(x, vec3(0, 0, 0)) + 0.5 \u0026lt;= feed_r: scene.set_voxel(feed_p + vec3(i, j, k), 2, feed_color) However, after completing Voxel Fortress, I felt that the whole thing wasn\u0026rsquo;t that complex; it could probably be done in around 50 lines of code.\nLet\u0026rsquo;s break down the entire pac-man: the surface is a sphere, the mouth is another internal sphere, the eyes are spheres, and the feed balls are also spheres. So, the main elements are all covered. How do we make the mouth open? Simply set the voxel of mat to 0, which effectively deletes it. Then, just insert a horizontal semi-cylinder into the mouth area, and it\u0026rsquo;s done!!\nVoxel Fortress Are there any voxel-based structures in reality? Besides LEGO, there are bricks! So, let\u0026rsquo;s build a fortress! (Maybe I\u0026rsquo;ll make a GW one later too!)\nBuilding this fortress is incredibly simple: just take a cube, add a top cap, and add raised brick blocks on all four sides. One function does it all!\n1 2 3 4 5 6 7 8 9 @ti.func def build_fortress(pos, sz1, sz2, height, color, color_noise): for x, y in ti.ndrange((-sz1, sz1 + 1), (-sz2, sz2 + 1)): if x == -sz1 or x == sz1 or y == -sz2 or y == sz2: for z in range(height): set_color_voxel(pos + vec3(x, z, y), 1, color, color_noise, 0.8) if (x + y) % 4 == 0 or (x + y) % 4 == 1: set_color_voxel(pos + vec3(x, height, y), 1, color, color_noise) set_color_voxel(pos + vec3(x, height - 2, y), 1, color, color_noise, 0.8) Then, we build one large central tower and four corner towers, and that\u0026rsquo;s it!\nNext, we construct the surrounding walls to enclose the fortress. A wall is essentially just a block; we can create a rectangle by specifying the bottom-left and top-right vertices and filling the middle. Of course, the walls should be slightly lower than the corner towers.\n1 2 3 4 5 6 @ti.func def build_block(pos1, pos2, color, color_noise, prob=1, mat=1): x_min, y_min, z_min = min(pos1.x, pos2.x), min(pos1.y, pos2.y), min(pos1.z, pos2.z) x_max, y_max, z_max = max(pos1.x, pos2.x), max(pos1.y, pos2.y), max(pos1.z, pos2.z) for x, y, z in ti.ndrange((x_min, x_max + 1), (y_min, y_max + 1), (z_min, z_max + 1)): set_color_voxel(vec3(x, y, z), mat, color, color_noise, prob) These walls look a bit fragile\u0026hellip; like they\u0026rsquo;d crumble at the slightest touch. Let\u0026rsquo;s make them thicker and leave space in the middle for soldiers to stand guard. But we also need to protect the soldiers from being hit, so we build them just like the fortress! Originally, the fortress faces were square; by changing them to rectangles and adding length and width parameters, we can create thick, sturdy walls!\nNow it really starts to feel like the real thing~\nLet\u0026rsquo;s build a gate; otherwise, how would anyone get in? The gate consists of just a sector and a rectangle, so easy~ The door frame is simply a larger version of the gate.\n1 2 3 4 for i in ti.ndrange((d_ - 2, d_ + 3)): build_door(vec3(0, 6, i), 6, 4, vec3(0.6, 0.6, 0.6), vec3(0)) build_door(vec3(0, 5, i), 5, 3, vec3(0, 0, 0), vec3(0), 1, 0) build_door(vec3(0, 5, d_), 5, 3, vec3(0.43, 0.352, 0.156), vec3(0)) Looks pretty good. Let\u0026rsquo;s add some ground: a layer of dirt topped with grass, and a path leading up to the gate. Also, let\u0026rsquo;s add windows to the towers on both sides. This is basically just deleting a door.\nInjecting some soul!! Adding little torches!!\n1 2 3 4 @ti.func def build_fire(pos): scene.set_voxel(pos, 2, vec3(1, 1, 0)) scene.set_voxel(pos + vec3(0, -1, 0), 1, vec3(0.43, 0.352, 0.156)) Finally, add a tiny doorknob, and we\u0026rsquo;re done!\nAnd last but not least, I optimized the lighting and exposure, and added a night mode.\nSome other thoughts Because I truly haven\u0026rsquo;t used Taichi in a long time. I previously watched the graphics course on Taichi on Bilibili, but unfortunately, I was busy with group meetings and various final-term courses back then, so I didn\u0026rsquo;t complete the major project at that time. I still feel a bit regretful about it. Fortunately, this time I have a chance to work on something I like and also get some hands-on practice! My code does have some areas that could be written better (mainly because I feel not yet fully proficient; many parts slightly conflict with common Python methods, so coding requires some mental shifts.\ndeep learning for pixel art\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-taichi-voxel-challenge-2022/","summary":"\u003ch2 id=\"voxel-pac-man\"\u003eVoxel Pac-man\u003c/h2\u003e\n\u003cp\u003eI came across \u003ccode\u003eTaichi\u003c/code\u003e\u0026rsquo;s \u003ca href=\"https://github.com/taichi-dev/voxel-challenge\"\u003e\u003ccode\u003evoxel-challenge\u003c/code\u003e\u003c/a\u003e repository quite a while ago. Since I\u0026rsquo;ve always been a huge fan of pixel art, I naturally kept an eye on voxel art as well; for instance, I previously conducted a \u003ccode\u003edeep learning for pixel art\u003c/code\u003e\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e survey on the topic. While casually browsing the repository, I stumbled upon an \u003ca href=\"https://github.com/taichi-dev/voxel-challenge/issues/1\"\u003eissue #1\u003c/a\u003e, which got me super excited to try it out myself. The code is available \u003ca href=\"https://github.com/yyyanbj/voxel-pac-man\"\u003ehere\u003c/a\u003e, and it roughly looks like this:\u003c/p\u003e","title":"Taichi Voxel Challenge 2022"},{"content":"Preface Actually, while working on hcaptcha challenge, I kept wondering: what will the next generation of CAPTCHA look like? What will the development path of CAPTCHA be?\nWhere is the Next Generation of CAPTCHA? The History of CAPTCHA? Human-Machine Confrontation? What Will the Next Generation of CAPTCHA Look Like?\nOrigins In the early days of the internet, many people used Leet Code1 to bypass keyword filtering.\nIn 2000, idrive.com began using CAPTCHA to protect its registration page; this was the first generation of CAPTCHA.\nBefore machine learning, especially deep learning, became widely adopted, the simplest way to handle this type of CAPTCHA was \u0026ldquo;crowdsourcing.\u0026rdquo; Images were sent to the server, which then distributed tasks to workers. This was the original \u0026ldquo;code-breaking platform.\u0026rdquo; Have you ever tried it? I encountered the term \u0026ldquo;code-breaking platform\u0026rdquo; back in middle school, and it might even have been my first pot of gold (?)—though I can\u0026rsquo;t remember for sure. Back then, the price per CAPTCHA was likely just a few cents. For an hour of work, you might earn about 1 yuan. Skilled workers could earn over 20 yuan a day. Some people treated it as typing practice hhh, and I approached it with the same mindset.\nLater, Optical Character Recognition (OCR) emerged to counter these CAPTCHAs. While this is very common now, in the early days when OCR technology was immature, few people attempted it. Defeating the early OCR was quite easy: just add some noise or a horizontal line, and it would fail. Consequently, the industry reverted to the \u0026ldquo;crowdsourcing\u0026rdquo; model.\nLater still, many image-based CAPTCHAs appeared in the form of math problems, requiring not only accurate character recognition but also an additional calculation step.\nDiverse Forms of CAPTCHA Nowadays, CAPTCHAs based on image character recognition have become less common after being cracked by OCR. They have been replaced by a wide variety of other CAPTCHA formats.\nFor example, click on the objects in the image in sequence. The characters may be tilted or distorted, and colors may vary. This requires OCR to have high robustness, and the data returned is simply click coordinates.\nThere are also CAPTCHAs that have evolved from text to icons or graphics. The underlying logic remains the same, but instead of text recognition, it has become an image matching problem.\nThen there are slider CAPTCHAs, where you simply drag the slider from left to right. The difficulty lies in the need to hold the slider down rather than just clicking the screen, and the sliding speed should not be constant.\nThere is also an improved version of the slider CAPTCHA: the puzzle CAPTCHA. For instance, GeeTest uses a puzzle CAPTCHA that adds a puzzle piece from the top image to the slider mechanism. There may be distractors, such as a completely unrelated puzzle piece suddenly appearing darkened in the image.\nImage Recognition Furthermore, we reached the era of advanced reCAPTCHA, which shifted the requirement to image recognition or object detection, asking you to select images containing buses or crosswalks. These CAPTCHAs have sparked a lot of complaints. Often, users feel they selected everything correctly but still fail verification. Sometimes, you are forced to click many times, and there are even response time limits. This has not only caused widespread dissatisfaction but also made many people question whether they are actually human hhhh.\nHowever, once a sufficient dataset is collected, the emergence of YOLOv5, with its high speed and lightweight performance, quickly became the nemesis of this type of CAPTCHA.\nhCAPTCHA also joined this battle and captured a significant market share. Recently, hCAPTCHA\u0026rsquo;s CAPTCHA upgrades truly caught my eye (e.g., vertical river, sky left airplant, elephant drawn with leaves). By incorporating Generative Adversarial Networks (GANs) into CAPTCHAs, they provide an infinite supply of data. But when I solved them using a simpler method ([1]2[2]3[3]4), I couldn\u0026rsquo;t help but question their very existence.\nThe Next Generation of CAPTCHA An interesting fact is that while CAPTCHA protects websites from attacks, the websites providing CAPTCHA are essentially running naked.\nCrawling the data from these \u0026ldquo;naked\u0026rdquo; CAPTCHAs is incredibly simple. For a 9-grid CAPTCHA, it is essentially a binary classification task because you only have two choices: click or don\u0026rsquo;t click. Moreover, due to the inherent limitations of CAPTCHA, high-resolution images cannot be used. This means the model can be as small as possible, and the image features will be very obvious.\nWhen an adversary keeps up with the frequency of CAPTCHA updates, it becomes terrifying. For every new task that appears, the adversary only needs less than one day to complete data labeling and model training.\nThis makes one wonder: if a CAPTCHA algorithm engineer spends over half a month creating a Generative Adversarial Network ready for production deployment, and an adversary completes data labeling and model training in less than a day, is it really worth it?\nFurthermore, the ultimate goal of CAPTCHA is to serve humans. If it degrades the user experience, is that a good outcome?\nRegarding the next generation of CAPTCHA, I would like to discuss it from two aspects:\nImage-based\nWill this kind of CAPTCHA have a future? (After all, everyone should be clear about how fiercely competitive the CV field in deep learning has become.) However, temporary use can still be effective. For instance, you could use results rendered in 3D to test a model\u0026rsquo;s 3D understanding capabilities. Currently, 3D understanding remains somewhat challenging, among other things.\nSimplify the Complex - CAPTCHA-less\nreCAPTCHA is already trying this, which is likely the mainstream direction for the next generation of CAPTCHAs. By monitoring the environment at the browser or system level to assess the probability of a malicious user, and combining this with user behavior to build a multimodal model, we can not only optimize the user experience (since users won\u0026rsquo;t even notice the CAPTCHA appearing) but also intercept malicious programs.\nReference Sources: 5.\nLeet Code\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[1]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[2]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[3]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nCAPTCHA - Wikipedia\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-where-is-the-next-generation-of-captcha/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eActually, while working on \u003ccode\u003ehcaptcha challenge\u003c/code\u003e, I kept wondering: what will the next generation of CAPTCHA look like? What will the development path of CAPTCHA be?\u003c/p\u003e\n\u003cp\u003eWhere is the Next Generation of CAPTCHA? The History of CAPTCHA? Human-Machine Confrontation? What Will the Next Generation of CAPTCHA Look Like?\u003c/p\u003e\n\u003ch2 id=\"origins\"\u003eOrigins\u003c/h2\u003e\n\u003cp\u003eIn the early days of the internet, many people used \u003ccode\u003eLeet Code\u003c/code\u003e\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e to bypass keyword filtering.\u003c/p\u003e\n\u003cp\u003eIn 2000, idrive.com began using CAPTCHA to protect its registration page; this was the first generation of CAPTCHA.\u003c/p\u003e","title":"What Will the Next Generation of CAPTCHA Look Like? A History of CAPTCHA and the Human-Machine Arms Race"},{"content":"It seems this prompt and the seaplane prompt have appeared with higher frequency recently; they must have been deployed to the production environment.\nSo naturally, I had to tackle it. The dataset can be found here .\nFirst, let\u0026rsquo;s analyze it. There are only three types in total: one made of leaves, one made of petals, and a third type made of some unknown metaphysical stuff that looks pitch black. Looking at the content composition, besides elephants, there are horses—it\u0026rsquo;s essentially a binary classification.\nActually, besides the prompt [\u0026quot;Please select all the elephants drawn with lеaves\u0026quot;], there is another similar one \u0026quot;Please select all the horses drawn with flowers\u0026quot;1, but I have almost never seen this prompt. I don\u0026rsquo;t know how the person who raised this issue managed to generate it. I think the bigger reason is that the perplexity of the flower images is too high, making them extremely hard to distinguish, which would greatly degrade the user experience. However, compared to humans extracting features, I feel machines can extract features much faster \u0026ndash;.\nThe approach is clear: first, the composition classification is as simple as it gets; we just need to extract the dominant color of the image. Here, I used kmeans for color clustering to extract the dominant color. I selected clustering centers k=3 to distinguish between bright and dark areas, so I assigned one center to each, leaving the third for the dominant color. This way, I only need to set a threshold to calculate the distance between the color and green. I chose 200, achieving a 100% discrimination accuracy.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 def _style_classification(self, img): # opencv to numpy img = np.array(img) # from 3d -\u0026gt; 2d img = img.reshape((img.shape[0] * img.shape[1], img.shape[2])).astype(np.float64) # print(img.shape) centroid, label = kmeans2(img, k=3) # print(centroid) # print(label) green_centroid = np.array([0.0, 255.0, 0.0]) flag = False min_dis = np.inf for i in range(len(centroid)): # distance between centroid and green \u0026lt; threshold # print(np.linalg.norm(centroid[i] - green_centroid)) min_dis = min(min_dis, np.linalg.norm(centroid[i] - green_centroid)) if min_dis \u0026lt; 200: flag = True return flag Distinguishing between elephants and horses is truly challenging, or maybe not challenging at all. Your image processing methods, such as further clustering or, like in the previous two blog posts [1]2[2]3, calculating weight or superpixel count, are quite difficult for this distinction. Since elephants and horses have similar volumes, and the ones below all have 5-6 support points (because elephants have trunks and tails, while horses have tails and mouths), it\u0026rsquo;s hard to tell them apart. To say it simply, isn\u0026rsquo;t this just \u0026quot;Dog vs Cat\u0026quot;? It\u0026rsquo;s merely an introductory image classification task in deep learning\u0026hellip;\nFirst, let\u0026rsquo;s label the data. We need to use that style of classifier to filter first, which will save us from labeling a lot of data.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 import os import sys sys.path.append(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) import cv2 from src.services.hcaptcha_challenger.solutions.elephant_solution import ElephantSolution img_path = os.path.join(\u0026#39;elephants_drawn_with_leaves\u0026#39;) label_file = open(\u0026#39;label.txt\u0026#39;, \u0026#39;w\u0026#39;) os.makedirs(os.path.join(img_path, \u0026#39;elephant\u0026#39;), exist_ok=True) os.makedirs(os.path.join(img_path, \u0026#39;house\u0026#39;), exist_ok=True) if __name__ == \u0026#39;__main__\u0026#39;: # 0 for house, 1 for elephant imgs = os.listdir(img_path) edwls = ElephantSolution() for idx, img_ in enumerate(imgs): img_path_ = os.path.join(img_path, img_) if os.path.isdir(img_path_): continue img = cv2.imread(img_path_) cv2.imshow(\u0026#34;img\u0026#34;, img) print(f\u0026#39;{img_path_}: {idx}\u0026#39;) if edwls._style_classification(img): key = cv2.waitKey(0) if key == ord(\u0026#39;0\u0026#39;): label_file.write(f\u0026#39;{img_path_} 0\\n\u0026#39;) label_file.flush() cv2.imwrite(os.path.join(img_path, \u0026#39;house\u0026#39;, img_), img) print(f\u0026#39;{img_path_} 0: house\u0026#39;) elif key == ord(\u0026#39;1\u0026#39;): label_file.write(f\u0026#39;{img_path_} 1\\n\u0026#39;) label_file.flush() cv2.imwrite(os.path.join(img_path, \u0026#39;elephant\u0026#39;, img_), img) print(f\u0026#39;{img_path_} 1: elephant\u0026#39;) else: print(\u0026#39;Drop\u0026#39;) Design a simple ResNet model. There\u0026rsquo;s no need for a massive model like ResNet18; this problem doesn\u0026rsquo;t even warrant it, so I DIYed a very small model, and I even downscaled the images to (64 x 64), which significantly reduces the number of parameters.\nTraining and testing in one go.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 import os import sys sys.path.append(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) import shutil import torch import torch.nn as nn import torch.nn.functional as F import torchvision import cv2 from PIL import Image from src.services.hcaptcha_challenger.solutions.elephant_solution import ElephantSolution class ResidualBlock(nn.Module): def __init__(self, in_channels, out_channels, stride=1): super(ResidualBlock, self).__init__() self.conv1 = nn.Conv2d(in_channels, out_channels, kernel_size=3, stride=stride, padding=1, bias=False) self.bn1 = nn.BatchNorm2d(out_channels) self.conv2 = nn.Conv2d(out_channels, out_channels, kernel_size=3, stride=1, padding=1, bias=False) self.bn2 = nn.BatchNorm2d(out_channels) self.relu = nn.ReLU(inplace=True) self.downsample = nn.Sequential() if stride != 1 or in_channels != out_channels: self.downsample = nn.Sequential( nn.Conv2d(in_channels, out_channels, kernel_size=1, stride=stride, bias=False), nn.BatchNorm2d(out_channels)) def forward(self, x): residual = x out = self.relu(self.bn1(self.conv1(x))) out = self.bn2(self.conv2(out)) out += self.downsample(residual) out = self.relu(out) return out class Net(nn.Module): def __init__(self, in_channels=3, num_classes=10): super(Net, self).__init__() self.conv1 = nn.Conv2d(in_channels, 16, kernel_size=7, stride=2, padding=3, bias=False) self.bn1 = nn.BatchNorm2d(16) self.relu = nn.ReLU(inplace=True) self.maxpool = nn.MaxPool2d(kernel_size=3, stride=2, padding=1) self.resblock1 = ResidualBlock(16, 32) self.resblock2 = ResidualBlock(32, 64, stride=2) self.avgpool = nn.AvgPool2d(kernel_size=7, stride=1) self.fc = nn.Linear(256, num_classes) def forward(self, x): x = self.conv1(x) x = self.bn1(x) x = self.relu(x) x = self.maxpool(x) x = self.resblock1(x) x = self.resblock2(x) x = self.avgpool(x) x = x.view(x.size(0), -1) # print(x.size()) x = self.fc(x) return x img_path = os.path.join(\u0026#39;..\u0026#39;, \u0026#39;database\u0026#39;, \u0026#39;elephants_drawn_with_leaves\u0026#39;) img_transform = torchvision.transforms.Compose([ # torchvision.transforms.Grayscale(num_output_channels=1), # torchvision.transforms.GaussianBlur(kernel_size=3), torchvision.transforms.Resize((64, 64)), torchvision.transforms.ToTensor(), ]) def train(): model = Net(3, 2) model.train() model.cuda() optimizer = torch.optim.Adam(model.parameters(), lr=0.005) # focal loss criterion = nn.CrossEntropyLoss() print(\u0026#39;model:\u0026#39;, model) data = torchvision.datasets.ImageFolder(img_path, transform=img_transform) data_loader = torch.utils.data.DataLoader(data, batch_size=1, shuffle=True) print(f\u0026#39;{len(data)} images\u0026#39;) epochs = 20 # train with focal loss for epoch in range(epochs): total_loss = 0 total_acc = 0 for i, (img, label) in enumerate(data_loader): img = img.cuda() label = label.cuda() optimizer.zero_grad() out = model(img) loss = criterion(out, label) loss.backward() optimizer.step() if (i + 1) % 10 == 0: print(f\u0026#39;epoch: {epoch + 1}, iter: {i + 1}, loss: {loss.item():.4f}\u0026#39;) total_loss += loss.item() total_acc += torch.sum(torch.argmax(out, dim=1) == label).item() print( f\u0026#39;epoch: {epoch + 1}, avg loss: {total_loss / len(data):.4f}, avg acc: {total_acc / len(data):.4f}\u0026#39; ) torch.save(model.state_dict(), \u0026#39;model.pth\u0026#39;) def test_single(model, img): img = img_transform(img) img = img.unsqueeze(0) img = img.cuda() out = model(img) pred = torch.argmax(out, dim=1) # print(f\u0026#39;pred: {pred.item()}\u0026#39;) if pred.item() == 0: return 0 else: return 1 def test(): model = Net(3, 2) model.load_state_dict(torch.load(\u0026#39;model.pth\u0026#39;)) model.eval() torch.onnx.export(model, torch.randn(1, 3, 64, 64), \u0026#39;model.onnx\u0026#39;, verbose=True, export_params=True) model.cuda() test_data_path = os.path.join(\u0026#39;val-dataset\u0026#39;) imgs = os.listdir(test_data_path) dir1 = os.path.join(\u0026#39;val-dataset\u0026#39;, \u0026#39;elephant_drawn_with_leaves\u0026#39;) dir2 = os.path.join(\u0026#39;val-dataset\u0026#39;, \u0026#39;house_drawn_with_leaves\u0026#39;) dir3 = os.path.join(\u0026#39;val-dataset\u0026#39;, \u0026#39;without_leaves\u0026#39;) dirs = [dir1, dir2, dir3] for dir in dirs: if os.path.exists(dir): shutil.rmtree(dir) os.mkdir(dir) es = ElephantSolution() for img in imgs: if os.path.isdir(os.path.join(test_data_path, img)): continue img_ = cv2.imread(os.path.join(test_data_path, img)) result = 2 if es._style_classification(img_): result = test_single(model, Image.open(os.path.join(test_data_path, img))) print(f\u0026#39;{img} is {result} save to {os.path.join(dirs[result], img)}\u0026#39;) cv2.imwrite(os.path.join(dirs[result], img), img_) if __name__ == \u0026#39;__main__\u0026#39;: train() test() Here is an interesting part: initially, after training, I tested it and found the error rate was quite high, around 10%, which is almost unacceptable because there are only 9 CAPTCHA images in total, making it tricky. I thought it was overfitting (this should be the conventional approach, after all, the training set accuracy was nearly 100%), so I increased the learning rate and decreased the number of epochs, but no matter what I did, the error rate stayed around 5%. I was baffled; how could such a simple classification task perform so poorly? Later\u0026hellip; I surprisingly increased the number of epochs, and found that the test set accuracy also reached nearly 100%. Finally, the overall error rate on the test set was around 1.8%, so I didn\u0026rsquo;t continue optimizing, and I was even too lazy to put the test data back into training.\nhCaptcha leaf elephant recognition demo (original GIF) View original GIF Finally, the entire model parameters were saved as pt, which is only 311KB. If exported as onnx, it becomes 290KB.\nHere I learned another trick: after exporting as onnx, I can directly use opencv for inference, which can greatly save resources in the deployment environment.\nSolution: The complete code is below\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 import os import sys sys.path.append(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) import cv2 import numpy as np from scipy.cluster.vq import kmeans2 class ElephantSolution: def __init__(self): self.debug = True def solution(self, img_stream, **kwargs) -\u0026gt; bool: # noqa \u0026#34;\u0026#34;\u0026#34;Implementation process of solution\u0026#34;\u0026#34;\u0026#34; img_arr = np.frombuffer(img_stream, np.uint8) img = cv2.imdecode(img_arr, flags=1) cv2.imshow(\u0026#34;img\u0026#34;, img) cv2.waitKey(0) if not self._style_classification(img): return False # using model to predict # print(img.shape) # resize img = cv2.resize(img, (64, 64)) # print(img.shape) model_path = os.path.join(\u0026#39;..\u0026#39;, \u0026#39;..\u0026#39;, \u0026#39;..\u0026#39;, \u0026#39;model\u0026#39;, \u0026#39;elephant_model.onnx\u0026#39;) model = cv2.dnn.readNetFromONNX(model_path) blob = cv2.dnn.blobFromImage(img, 1 / 255.0, (64, 64), (0, 0, 0), swapRB=True, crop=False) model.setInput(blob) out = model.forward() # print(out.shape) # print(out) label = np.argmax(out, axis=1)[0] # print(label) if label == 0: return True return False def _style_classification(self, img): # opencv to numpy img = np.array(img) # from 3d -\u0026gt; 2d img = img.reshape((img.shape[0] * img.shape[1], img.shape[2])).astype(np.float64) # print(img.shape) centroid, label = kmeans2(img, k=3) # print(centroid) # print(label) green_centroid = np.array([0.0, 255.0, 0.0]) flag = False min_dis = np.inf for i in range(len(centroid)): # distance between centroid and green \u0026lt; threshold # print(np.linalg.norm(centroid[i] - green_centroid)) min_dis = min(min_dis, np.linalg.norm(centroid[i] - green_centroid)) if min_dis \u0026lt; 200: flag = True return flag \u0026quot;Please select all the horses drawn with flowers\u0026quot;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[1]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[2]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-hcaptcha-elephant-drawn-with-leaves/","summary":"\u003cp\u003eIt seems this prompt and the seaplane prompt have appeared with higher frequency recently; they must have been deployed to the production environment.\u003c/p\u003e\n\u003cp\u003eSo naturally, I had to tackle it. The dataset can be found \u003ca href=\"https://github.com/QIN2DIM/img_pool/releases/tag/elephants_drawn_with_leaves\"\u003ehere \u003c/a\u003e.\u003c/p\u003e\n\n\u003cfigure class=\"align-center\"\u003e\u003ca href=\"/p/blog-hcaptcha-elephant-drawn-with-leaves/images/overview.png\"\u003e\n\u003cimg loading=\"lazy\" decoding=\"async\" src=\"/p/blog-hcaptcha-elephant-drawn-with-leaves/images/overview_hu16269298390059638712.webp\" alt=\"Overview of CAPTCHA samples for the leaf elephant task\" width=\"720\" height=\"384\" srcset=\"/p/blog-hcaptcha-elephant-drawn-with-leaves/images/overview_hu18158979934393581586.webp 360w, /p/blog-hcaptcha-elephant-drawn-with-leaves/images/overview_hu3181667444720115284.webp 480w, /p/blog-hcaptcha-elephant-drawn-with-leaves/images/overview_hu16269298390059638712.webp 720w, /p/blog-hcaptcha-elephant-drawn-with-leaves/images/overview_hu14903010596564395253.webp 1080w, /p/blog-hcaptcha-elephant-drawn-with-leaves/images/overview_hu10341424244034081744.webp 1440w\" sizes=\"(max-width: 748px) calc(100vw - 28px), 720px\"\u003e\n\u003c/a\u003e\n\u003c/figure\u003e\n\n\u003cp\u003eFirst, let\u0026rsquo;s analyze it. There are only three types in total: one made of leaves, one made of petals, and a third type made of some unknown metaphysical stuff that looks pitch black. Looking at the content composition, besides elephants, there are horses—it\u0026rsquo;s essentially a binary classification.\u003c/p\u003e","title":"hCaptcha Leaf Elephant Recognition: Data and Model Practice"},{"content":"Day 0 I participated in CodeCraft 2021 last year, but it conflicted with other contests and personal matters, so I didn\u0026rsquo;t even submit anything in the preliminary round and it ended there. This year, I thought I\u0026rsquo;d at least get a wristband, but who would\u0026rsquo;ve thought~~~ I got absolutely nothing.\nAbout the team name, I found it quite interesting. The name baseline was banned by the organizers. Last year, I used --baseline--. This year, having learned about many Unicode characters compared to last year, I thought I\u0026rsquo;d find a character to replace one, but who would\u0026rsquo;ve thought it only supports ASCII? I tried baseIine, but in the sans-serif font on that list, the letter \u0026lsquo;I\u0026rsquo; has no horizontal bar on top or bottom, making it look exactly like baseline. I guess the organizers never expected someone to name their team like this, hhhhh.\nPractice Contest The entire problem was relatively simple. There were some servers (edge nodes) and clients (client nodes). At each moment (abstracting a time period into a time node), clients had demands, and servers needed to allocate bandwidth to clients to meet those demands. However, the links between servers and clients had to satisfy latency requirements. Finally, the cost was calculated using the 95th percentile billing of the bandwidth sequence.\nAt first, I didn\u0026rsquo;t understand the 95th percentile billing correctly and thought it was the 95th percentile billing for each moment. I casually wrote a greedy algorithm and submitted it, surprisingly getting some points. The opening score was 1.34 million. I only realized my mistake when I wrote my own judger.\nAdditionally, when writing the greedy algorithm, I used pandas and finished it in under 20 minutes. After submitting, it kept running with errors. Later, I realized pandas wasn\u0026rsquo;t even available. I switched to numpy to read the txt file directly, splitting by commas. However, all the pandas indices I used in the code were completely broken, so I had no choice but to rewrite part of the code.\nAnother pitfall was that data reading had to start from the root directory. Ugh.\nThen, I thought it was just a load balancing problem, so I wrote an averaging method, which actually performed worse than the greedy approach? Then I thought it was an issue of selection order, so I tweaked several sorting methods, but they even performed worse than my first greedy submission (the first greedy submission was also sorted, by bandwidth and the number of client nodes it could handle, prioritizing servers with higher bandwidth and fewer client nodes).\nIn between, I considered binary search on the answer, setting a maximum threshold to see if it was feasible, and then trying to wrap it with a network flow algorithm.\nThis might not have been a load balancing problem at all. Then, I started from the 95th percentile billing and tried to \u0026lsquo;freeload\u0026rsquo; that 5%, and sure enough, the performance improved dramatically. I only needed to consider the order of demands, handling the 5% demand for each edge node separately. After some random tweaks, I got around 60w, and after adjusting some parameters, I reached 40w. Then, I just lay flat and submitted, ranking 2nd. I thought that would be enough to advance to the semi-finals, but in the last few days, my rank dropped from 2nd all the way to 20th. At that time, I hadn\u0026rsquo;t yet realized the severity of the problem - -.\nPreliminary Round I missed the first day of the preliminary round due to some matters. When I came back and submitted my previous best result, I found that the order of client nodes was no longer guaranteed. Ugh. After some adjustments and resubmitting, I only got around 40th place. I then tried a few of my own ideas, but the improvement was minimal. Ultimately, it was an algorithmic issue. I had already written network flow and minimum cost flow algorithms during the practice contest, but Python was too slow, and I really couldn\u0026rsquo;t compete. I felt that the approach of binary search combined with network flow was quite scientific.\nI gave up and focused on my graduation project instead, wasting time.\nFinal Thoughts To be honest, although I was lazy, the overall experience of participating was quite poor. During the practice contest, it was fine; everyone was asking normal questions about submissions. But later, people started exchanging ideas, acting like they were dividing apples among themselves? In the last few days of the practice contest, everyone suddenly became smart, hehe. In the preliminary round, even more smart people appeared, hehe. Many people I had never seen before in the practice contest showed up, hehe. Truly impressive.\nPostscript I made it to the top 64 in the Guangdong-Hong Kong-Macao Greater Bay Area and got a certificate, hhhh.\nI won\u0026rsquo;t participate next year.\n","permalink":"https://blog.bj-yan.top/en/p/blog-codecraft-2022/","summary":"\u003ch2 id=\"day-0\"\u003eDay 0\u003c/h2\u003e\n\u003cp\u003eI participated in CodeCraft 2021 last year, but it conflicted with other contests and personal matters, so I didn\u0026rsquo;t even submit anything in the preliminary round and it ended there. This year, I thought I\u0026rsquo;d at least get a wristband, but who would\u0026rsquo;ve thought~~~ I got absolutely nothing.\u003c/p\u003e\n\u003cp\u003eAbout the team name, I found it quite interesting. The name \u003ccode\u003ebaseline\u003c/code\u003e was banned by the organizers. Last year, I used \u003ccode\u003e--baseline--\u003c/code\u003e. This year, having learned about many Unicode characters compared to last year, I thought I\u0026rsquo;d find a character to replace one, but who would\u0026rsquo;ve thought it only supports ASCII? I tried \u003ccode\u003ebaseIine\u003c/code\u003e, but in the sans-serif font on that list, the letter \u0026lsquo;I\u0026rsquo; has no horizontal bar on top or bottom, making it look exactly like \u003ccode\u003ebaseline\u003c/code\u003e. I guess the organizers never expected someone to name their team like this, hhhhh.\u003c/p\u003e","title":"The Helplessness of baseIine - My Journey in CodeCraft 2022"},{"content":" My machine is a living life. I\u0026rsquo;ll prove it.\nIntroduction hCaptcha has just received another update today, introducing a new tag that requires selecting in the sky left-flying airplanes.\nFortunately, however, every sample image contains an airplane, so this tag effectively removes one constraint, leaving only the two constraints of in the sky and flying left.\nMain Method First, let\u0026rsquo;s observe the images, still referring to the collected dataset1. Besides the fact mentioned earlier that each sample image contains an airplane, if the airplane is in the sky, its background must be very \u0026ldquo;clean\u0026rdquo;. If it is not in the sky, it can basically be judged as being on the ground. Images on the ground also consist of multiple regions, such as lawns, airport runways, background forests, background skies, etc.\nSo, the key to distinguishing the first problem is the complexity of the background. How to do it?\nMy initial idea was to inherit the previous approach of color block filtering, but after dividing the color blocks, it was difficult to distinguish whether a block was an airplane or a background block, so it was discarded.\nThen, I looked at most methods for removing sky backgrounds, which are basically based on threshold filtering in the HSV color space, followed by morphological operations like erosion and dilation for noise reduction. I tested it a few times myself, but the color range of the sky is slightly too large, containing blue, white, and yellow. More critically, it is very similar to the color of airplanes, as airplanes are mostly light colors like pale blue or white. This was also discarded.\nI then thought that the shadow under the airplane would be black, so creating a superpixel smart selection based on black could work, but I don\u0026rsquo;t know how to implement it (x, so it was discarded.\nFinally, the adopted solution was to use the contour line method Canny to find all contour lines, then set a threshold to determine if the airplane is in the sky based on the number of contour lines. If it is in the sky, the contour lines will be very simple, whereas if it is not in the sky, a large amount of chaotic lines will be added. Of course, this threshold was derived by randomly testing a few images, given the huge difference between the two cases. After such processing, the judgment accuracy approaches 100%.\nOkay, one problem solved. Now, the remaining problem is: how to determine if the airplane is facing left?\nThis really stumped me. Without using Deep Learning, it is indeed difficult, but there are some clever tricks. First, most airplanes are transport planes, fighters, or passenger jets; there are rarely propeller planes in the front, and I haven\u0026rsquo;t seen helicopters either. There are even WTF Airplanes like the one below (what the heck is this?)\nSince that\u0026rsquo;s the case, these types of airplanes have a characteristic: \u0026ldquo;light head, heavy tail\u0026rdquo;. Besides the tail fin being heavier and the nose being pointed and lighter, the wings also point backward, so the \u0026ldquo;center of gravity\u0026rdquo; of the entire image should be biased towards the tail. I could determine the airplane\u0026rsquo;s direction based on 4 points: extreme left, extreme right, midpoint, and center of gravity.\nSounds scientific, right? However, the actual effect was not very good. The most critical issue is that airplanes have perspective relationships, so from the front view, the center of gravity might appear to be at the front. For example, if the nose is facing you, a large number of lines are drawn above the nose, while there are very few lines at the tail, making the center of gravity of the entire image appear at the front. Later, I wondered if I could fill the image, but the resulting contour lines were mostly not closed, making it difficult to fill the entire airplane (if that were possible, image segmentation would be simple).\nLater, I had a sudden inspiration and took a different path. Still observing the contour lines drawn above, the nose, lacking complex elements, produces relatively \u0026ldquo;simple\u0026rdquo; contour lines, while the tail, due to components like \u0026ldquo;tail fins\u0026rdquo;, produces relatively \u0026ldquo;complex\u0026rdquo; ones. So how to measure this \u0026ldquo;simplicity\u0026rdquo; and \u0026ldquo;complexity\u0026rdquo;? Just sum them up\u0026hellip; That\u0026rsquo;s right, in the end, I counted the non-zero pixels from x_min to x_min + left_threshold on the left (which are the contour line pixels) and compared them with the pixel count from x_max - left_threshold to x_max. Whichever is larger is the tail; if the right side is larger, then the nose is on the left. Unexpectedly, the final result was quite good, basically passing verification within 1-2 rounds.\nhCaptcha airplane direction recognition demo (original GIF) View original GIF Attached is the complete code for the test version\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 from itertools import count import cv2 import numpy as np import matplotlib.pyplot as plt from scipy import ndimage as ndi from skimage.util import random_noise from skimage import feature class SkyLeftAirplaneChallenger: \u0026#34;\u0026#34;\u0026#34;A fast solution for identifying vertical rivers\u0026#34;\u0026#34;\u0026#34; def __init__(self): self.flag = \u0026#34;skyleftairplane_model\u0026#34; self.sky_threshold = 1800 self.left_threshold = 30 self.debug = True @staticmethod def _remove_border(img): img[:, 1] = 0 img[:, -2] = 0 img[1, :] = 0 img[-2, :] = 0 return img def solution(self, img_stream, **kwargs) -\u0026gt; bool: # noqa \u0026#34;\u0026#34;\u0026#34;Implementation process of solution\u0026#34;\u0026#34;\u0026#34; img_arr = np.frombuffer(img_stream, np.uint8) img = cv2.imdecode(img_arr, flags=1) img = cv2.cvtColor(img, cv2.COLOR_BGR2GRAY) # cv2.imshow(\u0026#34;img\u0026#34;, img) # cv2.waitKey(0) edges1 = feature.canny(img) edges1 = self._remove_border(edges1) edges2 = feature.canny(img, sigma=3) edges2 = self._remove_border(edges2) # display results # fig, ax = plt.subplots(nrows=1, ncols=3, figsize=(8, 3)) # ax[0].imshow(img, cmap=\u0026#39;gray\u0026#39;) # ax[0].set_title(\u0026#39;noisy image\u0026#39;, fontsize=20) # ax[1].imshow(edges1, cmap=\u0026#39;gray\u0026#39;) # ax[1].set_title(r\u0026#39;Canny filter, $\\sigma=1$\u0026#39;, fontsize=20) # ax[2].imshow(edges2, cmap=\u0026#39;gray\u0026#39;) # ax[2].set_title(r\u0026#39;Canny filter, $\\sigma=3$\u0026#39;, fontsize=20) # for a in ax: # a.axis(\u0026#39;off\u0026#39;) # fig.tight_layout() # plt.show() # fill_plane = ndi.binary_fill_holes(edges1) # fig, ax = plt.subplots(figsize=(4, 3)) # ax.imshow(fill_plane, cmap=plt.cm.gray) # ax.set_title(\u0026#39;filling the holes\u0026#39;) # ax.axis(\u0026#39;off\u0026#39;) # plt.show() # print(np.count_nonzero(edges1)) # print(np.count_nonzero(edges2)) if np.count_nonzero(edges1) \u0026gt; self.sky_threshold: if self.debug: print(\u0026#39;[not in sky] \u0026#39;, end=\u0026#39;\u0026#39;) return False # get avg coordinate of edges where edges are not zero # avg_point = np.average(np.nonzero(edges1), axis=1) # print(avg_point) min_x = np.min(np.nonzero(edges1), axis=1)[1] max_x = np.max(np.nonzero(edges1), axis=1)[1] left_nonzero = np.count_nonzero(edges1[:, min_x:min(max_x, min_x + self.left_threshold)]) right_nonzero = np.count_nonzero(edges1[:, max(min_x, max_x - self.left_threshold):max_x]) # print(left_nonzero, right_nonzero) if left_nonzero \u0026gt; right_nonzero: if self.debug: print(\u0026#39;[not turn left] \u0026#39;, end=\u0026#39;\u0026#39;) return False # mid_x = (min_x + max_x) / 2 # print(min_x, max_x, mid_x, avg_point[0] \u0026lt; mid_x) # if avg_point[0] \u0026gt;= mid_x: # return False # plt.show() return True if __name__ == \u0026#39;__main__\u0026#39;: import os import sys sys.path.append(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) result_path = \u0026#39;result.txt\u0026#39; if os.path.exists(result_path): os.remove(result_path) # result_file = open(result_path, \u0026#39;w\u0026#39;) result_file = sys.stdout base_path = os.path.join(\u0026#39;..\u0026#39;, \u0026#39;database\u0026#39;, \u0026#39;airplane_in_the_sky_flying_left\u0026#39;) image_list = os.listdir(base_path) # image_list.sort() for image_name in image_list: image_path = os.path.join(base_path, image_name) with open(image_path, \u0026#34;rb\u0026#34;) as file: data = file.read() solution = SkyLeftAirplaneChallenger().solution(data) result_file.write(f\u0026#39;{image_name}: {solution}\\n\u0026#39;) result_file.flush() result_file.close() Conclusion To be honest, solving a high-quality image processing problem from hCaptcha every day still feels pretty cool hhhh.\nHowever, through sharing by netizens, I have seen more problems solved using generative models. After all, they are paid to do this, and later some tasks can no longer be solved just by image processing, such as black-and-white striped cats that appeared in feedback for certain Tampermonkey plugins.\nJust a clever workaround to meet each challenge.\ndataset\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-hcaptcha-sky-left-airplane/","summary":"\u003cblockquote\u003e\n\u003cp\u003eMy machine is a living life. I\u0026rsquo;ll prove it.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\n\u003cfigure class=\"align-center\"\u003e\u003ca href=\"/p/blog-hcaptcha-sky-left-airplane/images/example.jpg\"\u003e\n\u003cimg loading=\"lazy\" decoding=\"async\" src=\"/p/blog-hcaptcha-sky-left-airplane/images/example_hu8139913857821317050.webp\" alt=\"CAPTCHA samples for the left-flying airplane task\" width=\"589\" height=\"786\" srcset=\"/p/blog-hcaptcha-sky-left-airplane/images/example_hu5615645687526737917.webp 360w, /p/blog-hcaptcha-sky-left-airplane/images/example_hu3147666006069205416.webp 480w, /p/blog-hcaptcha-sky-left-airplane/images/example_hu8139913857821317050.webp 589w\" sizes=\"(max-width: 617px) calc(100vw - 28px), 589px\"\u003e\n\u003c/a\u003e\n\u003c/figure\u003e\n\n\u003cp\u003e\u003ca href=\"https://www.hcaptcha.com/\"\u003ehCaptcha\u003c/a\u003e has just received another update today, introducing a new tag that requires selecting \u003cstrong\u003ein the sky\u003c/strong\u003e \u003cstrong\u003eleft-flying\u003c/strong\u003e \u003cstrong\u003eairplanes\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eFortunately, however, every sample image contains an airplane, so this tag effectively removes one constraint, leaving only the two constraints of \u003cstrong\u003ein the sky\u003c/strong\u003e and \u003cstrong\u003eflying left\u003c/strong\u003e.\u003c/p\u003e","title":"hCaptcha Left-Flying Airplane Recognition: Image Processing Approach"},{"content":" Introduction Recently, hCaptcha1 underwent an update, shifting from previous object recognition to selecting vertical rivers within images. As shown in the image above, screenshot from 1.\nTo be honest, the first glance revealed it was done by some generative models because the image features are obvious: some areas are particularly blurry, and the generation logic is quite nonsensical. Initially, I thought it was an application of GPT-3, converting image descriptions into text, but later deemed it unreasonable due to the lack of imagination. Moreover, the distribution of elements is extremely obvious; everything is straight and unnatural. It must be some work on generating raw images from semantic images, such as NVIDIA\u0026rsquo;s artistic creations.\nBefore starting, I never imagined that NVIDIA\u0026rsquo;s GANs could be applied to the CAPTCHA field, nor did I expect that after the upgrade, it would be downgraded from object detection to requiring only image processing to pass.\nMain Content @QIN2DIM previously wrote about a hCaptcha-challenger2 project using YOLOv5 to handle this. After the update, they released a dataset vertical_river3, so I used image processing to quickly implement a Demo.\nFirst, the image features are very obvious, mainly consisting of several elements: sky, mountains, grass, and water in the background, with very distinct divisions and distributions, all appearing as separate blocks.\nMy initial idea was to apply some filters, such as 高斯滤波, and then directly use slic for superpixel segmentation. The expected result was that superpixels would group each color block into a single superpixel. However, the final result was not ideal because superpixel segmentation relies heavily on factors other than color, and limiting the number of superpixels causes larger color ranges to be grouped into one class. Here is an example.\nAs you can see, there are many issues: some parts that should be segmented were not correctly divided and instead clustered together, while some parts that shouldn\u0026rsquo;t be segmented were clustered together due to the filtering effect.\nAfter repeatedly adjusting parameters with no success, I felt that this segmentation approach should not treat a single superpixel as a color block, but rather use the superpixel segmentation results as a reference or boundary for segmentation.\nI originally wanted to try some edge enhancement algorithms but couldn\u0026rsquo;t find a suitable one because I observed that most CAPTCHA images have very blurry boundaries. Therefore, preserving edge information during segmentation is crucial. When selecting filters, it is essential to prioritize whether the filter can correctly preserve edge information rather than blurring the edges. Among them, 双边滤波 is an excellent choice, and 均值偏移 is also a good option.\nAfter these two steps, you have become myopic, but the edges remain relatively clear. Now you can proceed with color block segmentation. Here, I mainly used scikit-image.graph.rag_mean_color to obtain average color blocks, primarily referencing the official RAG Merge4 implementation.\nThe judgment condition was written simply: I only checked if the last row contains three or more color blocks. If so, I consider that the middle is separated by a river, indicating the presence of a vertical river. I feel this judgment condition could still be optimized, but after testing, out of about 100 images, only around 2 were misclassified. Even without optimization, this accuracy is acceptable.\nThe final effect is as follows:\nI submitted a PR, and the effect after @QIN2DIM merged it is as follows:\nhCaptcha vertical river recognition demonstration (original GIF) View original GIF Code used for testing and visualization:\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 #!/usr/bin/env python # -*- coding: utf-8 -*- # @File : src\\services\\hcaptcha_challenger\\river_challenger.py # @Time : 2022-03-01 20:32:08 # @Author : Bingjie Yan # @Email : bj.yan.pa@qq.com # @License : Apache License 2.0 import cv2 import numpy as np import skimage from skimage.morphology import disk from skimage.segmentation import watershed, slic, mark_boundaries from skimage.filters import rank from skimage.util import img_as_ubyte from skimage.color import rgb2gray, label2rgb from skimage.future import graph from scipy import ndimage as ndi import matplotlib.pyplot as plt def _weight_mean_color(graph, src, dst, n): \u0026#34;\u0026#34;\u0026#34;Callback to handle merging nodes by recomputing mean color. The method expects that the mean color of `dst` is already computed. Parameters ---------- graph : RAG The graph under consideration. src, dst : int The vertices in `graph` to be merged. n : int A neighbor of `src` or `dst` or both. Returns ------- data : dict A dictionary with the `\u0026#34;weight\u0026#34;` attribute set as the absolute difference of the mean color between node `dst` and `n`. \u0026#34;\u0026#34;\u0026#34; diff = graph.nodes[dst][\u0026#39;mean color\u0026#39;] - graph.nodes[n][\u0026#39;mean color\u0026#39;] diff = np.linalg.norm(diff) return {\u0026#39;weight\u0026#39;: diff} def merge_mean_color(graph, src, dst): \u0026#34;\u0026#34;\u0026#34;Callback called before merging two nodes of a mean color distance graph. This method computes the mean color of `dst`. Parameters ---------- graph : RAG The graph under consideration. src, dst : int The vertices in `graph` to be merged. \u0026#34;\u0026#34;\u0026#34; graph.nodes[dst][\u0026#39;total color\u0026#39;] += graph.nodes[src][\u0026#39;total color\u0026#39;] graph.nodes[dst][\u0026#39;pixel count\u0026#39;] += graph.nodes[src][\u0026#39;pixel count\u0026#39;] graph.nodes[dst][\u0026#39;mean color\u0026#39;] = (graph.nodes[dst][\u0026#39;total color\u0026#39;] / graph.nodes[dst][\u0026#39;pixel count\u0026#39;]) class RiverChallenger(object): def __init__(self) -\u0026gt; None: pass def challenge(self, img_stream): img_arr = np.frombuffer(img_stream, np.uint8) img = cv2.imdecode(img_arr, flags=1) height, width = img.shape[:2] # # filter img = cv2.pyrMeanShiftFiltering(img, sp=10, sr=40) img = cv2.bilateralFilter(img, d=9, sigmaColor=100, sigmaSpace=75) # # enhance brightness # img_hsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV) # img_hsv[:, :, 2] = img_hsv[:, :, 2] * 1.2 # img = cv2.cvtColor(img_hsv, cv2.COLOR_HSV2BGR) labels = slic(img, compactness=30, n_segments=400, start_label=1) g = graph.rag_mean_color(img, labels) labels2 = graph.merge_hierarchical(labels, g, thresh=35, rag_copy=False, in_place_merge=True, merge_func=merge_mean_color, weight_func=_weight_mean_color) # view results out = label2rgb(labels2, img, kind=\u0026#39;avg\u0026#39;, bg_label=0) out = mark_boundaries(out, labels2, (0, 0, 0)) skimage.io.imshow(out) skimage.io.show() print(np.unique(labels2[-1])) ref_value = len(np.unique(labels2[-1])) return ref_value \u0026gt;= 3 if __name__ == \u0026#39;__main__\u0026#39;: import os import sys sys.path.append(os.path.dirname(os.path.dirname(os.path.abspath(__file__)))) # result_path = \u0026#39;result.txt\u0026#39; # if os.path.exists(result_path): # os.remove(result_path) # result_file = open(result_path, \u0026#39;w\u0026#39;) result_file = sys.stdout base_path = os.path.join(\u0026#39;database\u0026#39;, \u0026#39;_challenge\u0026#39;) list_dirs = os.listdir(base_path) for dir in list_dirs[:1]: print(dir, file=result_file) for i in range(1, 10): img_filepath = os.path.join(base_path, dir, f\u0026#39;挑战图片{i}.png\u0026#39;) with open(img_filepath, \u0026#34;rb\u0026#34;) as file: data = file.read() rc = RiverChallenger() result = rc.challenge(data) print(f\u0026#39;挑战图片{i}.png:{result}\u0026#39;, file=result_file) Conclusion Neural networks have made everything more complex, yet also simpler.\n(Discussing the experience of a hCaptcha engineer seeing their project, developed over half a day, surpassed by a much simpler algorithm in less than a night?)\nhCaptcha\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhCaptcha-challenger\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nvertical_river\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRAG Merge\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-hcaptcha-vertical-river/","summary":"\u003cfigure class=\"align-center\"\u003e\u003ca href=\"/p/blog-hcaptcha-vertical-river/images/demo.jpg\"\u003e\n\u003cimg loading=\"lazy\" decoding=\"async\" src=\"/p/blog-hcaptcha-vertical-river/images/demo_hu9723212889169720686.webp\" alt=\"CAPTCHA samples for the vertical river task\" width=\"593\" height=\"762\" srcset=\"/p/blog-hcaptcha-vertical-river/images/demo_hu7561029871307245844.webp 360w, /p/blog-hcaptcha-vertical-river/images/demo_hu14306872510471133959.webp 480w, /p/blog-hcaptcha-vertical-river/images/demo_hu9723212889169720686.webp 593w\" sizes=\"(max-width: 621px) calc(100vw - 28px), 593px\"\u003e\n\u003c/a\u003e\n\u003c/figure\u003e\n\n\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eRecently, \u003ccode\u003ehCaptcha\u003c/code\u003e\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e underwent an update, shifting from previous object recognition to selecting vertical rivers within images. As shown in the image above, screenshot from \u003csup id=\"fnref1:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eTo be honest, the first glance revealed it was done by some generative models because the image features are obvious: some areas are particularly blurry, and the generation logic is quite nonsensical. Initially, I thought it was an application of GPT-3, converting image descriptions into text, but later deemed it unreasonable due to the lack of imagination. Moreover, the distribution of elements is extremely obvious; everything is straight and unnatural. It must be some work on generating raw images from semantic images, such as NVIDIA\u0026rsquo;s artistic creations.\u003c/p\u003e","title":"hCaptcha Vertical River Recognition: Image Features and Implementation Approach"},{"content":"Introduction I have actually included hyperlinks to my personal resume on multiple websites, all pointing to the resume on my homepage. However, updating that homepage resume has become quite troublesome due to the frequent need for modifications and copying, so I haven\u0026rsquo;t updated it in a long time. After nearly a year of working with CI/CD workflows, I\u0026rsquo;ve gained a solid grasp of GitHub Actions, which led me to develop this resume compilation and distribution workflow. I\u0026rsquo;m documenting the pitfalls I\u0026rsquo;ve encountered along the way.\nWorkflow Actually, the entire workflow is quite simple.\nCompile the tex file in the repository into a pdf file. Place the compiled pdf file into the pdf directory of my personal homepage and push the changes. (Sounds simple, but I encountered quite a few pitfalls, spending nearly 6 hours on it.)\nThe compilation part didn\u0026rsquo;t take long. Initially, I wanted to use latexmk-actions1, but after checking, it didn\u0026rsquo;t fully meet my needs, and since it was still under active development, I abandoned it. I eventually switched to xu-cheng/latex-action@v22. I have to say, the author\u0026rsquo;s Actions configuration is excellent; it basically fulfilled all my requirements, including handling multiple directories and files without issue.\nThe main pitfall I encountered was related to 鉴权.\nSince I had previously handled cross-repository operations, I typically just used an action like hugo-deploy or gh-pages-deploy, configuring the secrets accordingly. Naturally, I assumed I could handle the push operation with a simple action this time as well. I looked at many Actions related to push; the best-written one was ad-m/github-push-action@master3, but I kept getting the error Error: Invalid exit code: 128. I checked several similar issues reported by the author, but they didn\u0026rsquo;t resolve it. Google yielded no results, so I had to look for other solutions. I examined a few other options, but they also failed to solve my problem.\nGoing back to basics, we only need authentication plus the push operation. I wondered if I could authenticate during a Pull or a Push. I tried using the https://[username]:${{ secrets.PUBLISH_KEY }}@github.com/[username]/[repo] format during a Push, but it didn\u0026rsquo;t work. It didn\u0026rsquo;t work during a Pull either (though later I thought this approach might actually be feasible: just clone the repo, remove the .git information, treat it as a new repository, then add remote url, add that link, perform the Push, and include the --force parameter. It should work, but I don\u0026rsquo;t know if anyone has tried it hhh).\nActually, the final solution was the simplest method: utilizing actions/checkout instead of cloning. The official checkout action provides the ssh-key parameter (note: this is not token, unless you are using Personal Access Token instead of Deploy Key below).\nThe final workflow is as follows. You can view the file version here4.\n1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 name: LaTex Compile and Push on: push: branches: [main] jobs: build: runs-on: ubuntu-latest steps: - name: Set up Git repository uses: actions/checkout@v2 with: persist-credentials: false - name: Set up HomePage uses: actions/checkout@v2 with: repository: yyyanbj/yyyanbj.github.io ref: master path: homepage ssh-key: ${{ secrets.PUBLISH_KEY }} - name: Compile CV uses: xu-cheng/latex-action@v2 with: root_file: | cv_en/cv_en.tex cv_cn/cv_cn.tex work_in_root_file_dir: true latexmk_use_xelatex: true - name: Setup repository run: | ls -l ${GITHUB_WORKSPACE} ls -l ${GITHUB_WORKSPACE}/cv_cn ls -l ${GITHUB_WORKSPACE}/cv_en git config --global user.email \u0026#34;41898282+github-actions[bot]@users.noreply.github.com\u0026#34; git config --global user.name \u0026#34;github-actions[bot]\u0026#34; cp ${GITHUB_WORKSPACE}/cv_cn/cv_cn.pdf homepage/pdf/cv_cn.pdf cp ${GITHUB_WORKSPACE}/cv_en/cv_en.pdf homepage/pdf/cv_en.pdf cd homepage \u0026amp;\u0026amp; git commit -am \u0026#34;Update CV\u0026#34; git status git push origin master Implementation This resume directory has two versions: one in Chinese and one in English. All required font resource files have been uploaded, so you can clone it directly and use it locally. The template was mainly modified from this project5, with some optimizations for Chinese characters.\nAdditionally, if you have a personal homepage and lack a LaTeX environment but want to compile or distribute the resume in the cloud, you can simply fork this repository, modify the cv_en.tex and cv_cn.tex files, and configure the workflows. The following section primarily covers the execution for those with such requirements.\nThe following content includes the following declarations: 主仓库: The repository where the PDF needs to be uploaded, e.g., yyyanbj/yyyanbj.github.io 本仓库: The repository containing the resume tex file, e.g., yyyanbj/cv\nCreate the Resume Open [https://github.com/yyyanbj/cv]，点击右上角 to fork, navigate to the forked repository under your username, and clone it.\nModify \u0026amp; Compile the Resume Modify the tex file according to your needs, as well as the tex file within the subfolder cv.\nDistribute Generate Deploy Key, refer to GitHub Authentication6\nYou can follow these steps: 1. ssh-keygen -t rsa -b 4096 -C \u0026quot;your_email@example.com\u0026quot; 2. Add id_rsa.pub to Settings -\u0026gt; Deploy Keys in the main repository; the name can be arbitrary, but you must grant write permissions 3. Add id_rsa to Settings -\u0026gt; Secrets in this repository, and fill in the name as PUBLISH_KEY\nAllow GitHub Actions to run, modify the content in .github/workflows/latex.yml as needed (such as repository name, branch name, etc.), then push the changes to the GitHub repository and wait for the build to run.\nConclusion I hope this helps. If this article or repository infringes upon any of your rights, please contact bj.yan.pa@qq.com, and I will handle it as soon as possible.\nlatexmk-actions\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nxu-cheng/latex-action@v2\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nad-m/github-push-action@master\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhere\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nthis project\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nGitHub Authentication\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-github-actions-cv/","summary":"\u003ch2 id=\"introduction\"\u003eIntroduction\u003c/h2\u003e\n\u003cp\u003eI have actually included hyperlinks to my personal resume on multiple websites, all pointing to the resume on my homepage. However, updating that homepage resume has become quite troublesome due to the frequent need for modifications and copying, so I haven\u0026rsquo;t updated it in a long time. After nearly a year of working with CI/CD workflows, I\u0026rsquo;ve gained a solid grasp of GitHub Actions, which led me to develop this resume compilation and distribution workflow. I\u0026rsquo;m documenting the pitfalls I\u0026rsquo;ve encountered along the way.\u003c/p\u003e","title":"Best Practices for Compiling and Distributing Resumes with GitHub Actions"},{"content":"Trading Platforms OpenSea: https://opensea.io\nThe center of the universe: can issue public chains (requires Gas Fee) and sidechains like Polygon (no Gas Fee)\nRarible: https://rarity.io\nCan issue public chains; charges transaction fees and Gas Fee after trades\nBigverse: https://www.nftcn.com.cn/\nDomestic sidechain\nTheOne.art: https://theone.art/\nSome say it\u0026rsquo;s on Polygon, but I haven\u0026rsquo;t used it myself; needs verification\nNeeds further investigation\nRedCave: https://www.redcave.com/\nStill in beta; not sure which chain it\u0026rsquo;s on\nIMO Blue Cat Digital: https://www.lanmsz.cn/\nConsortium chain. Not sure about third-party integrations. They claim not to rely on any platform but use a consortium chain hhh. However, this seems to be combined with game IPs; let\u0026rsquo;s see how their world-building develops.\nThe following are all private chains from big tech companies with no collectible value\nJingtan (Alipay AntChain)\nHuanhe (Tencent)\nLingxi (JD.com)\nXihuang (Baidu)\nDongyiyuandian (Baidu)\nTianxia Hong Universe (Xianyu)\nR-SPACE (Xiaohongshu)\nNetEase Planet (NetEase)\nYouBanquan\nIBox\nShuangjing Museum\nHechao Wenguan\nCryptoSpace\nIntangible Cultural Heritage Digital Collection Platform\nMineNFT Youyu Block\nOne Meta\nNews \u0026amp; Updates rarity.tools: https://rarity.tools/upcoming/\nNew avatar NFT drops\nNFTCALENDAR: https://nftcalendar.io/\nNFT drop schedule\nGalleries Essentially drives traffic to trading platforms\nNiftyGateWay: https://niftygateway.com/\nHas high requirements for creators; quality assurance is good\nartblock: https://www.artblock.io/gallery\nTools \u0026amp; Others NFT Artist Assistant: https://www.uonus.net/\nAn open-source small tool. The algorithm is just so-so; honestly, it\u0026rsquo;s a minor side project.\nnft-image-generator : https://github.com/benyaminahmed/nft-image-generator\nThe tool written by the kid who drew the whale using Jupyter Notebook is only a few lines long.\nnft generator: https://github.com/cyberdoggos/generator\nA tool for automatically generating NFTs, mainly by arranging and combining layers, supporting rarity levels.\nsolseum-nft-generator: https://github.com/Solseum/solseum-nft-generator\nNFT_Art_Generator: https://github.com/kosmosmo/NFT_Art_Generator\nSpongeBob NFT\nnft-ganesha: https://github.com/ekkyarmandi/nft-ganesha\nGanesha NFT\nConclusion Alright, that\u0026rsquo;s about all the content. Feel free to leave me a message if you have any questions hhh\n","permalink":"https://blog.bj-yan.top/en/p/blog-nft-common-site/","summary":"\u003ch2 id=\"trading-platforms\"\u003eTrading Platforms\u003c/h2\u003e\n\u003cp\u003eOpenSea: \u003ca href=\"https://opensea.io\"\u003ehttps://opensea.io\u003c/a\u003e\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eThe center of the universe: can issue public chains (requires Gas Fee) and sidechains like Polygon (no Gas Fee)\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eRarible: \u003ca href=\"https://rarity.io\"\u003ehttps://rarity.io\u003c/a\u003e\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eCan issue public chains; charges transaction fees and Gas Fee after trades\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eBigverse: \u003ca href=\"https://www.nftcn.com.cn/\"\u003ehttps://www.nftcn.com.cn/\u003c/a\u003e\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eDomestic sidechain\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eTheOne.art: \u003ca href=\"https://theone.art/\"\u003ehttps://theone.art/\u003c/a\u003e\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eSome say it\u0026rsquo;s on Polygon, but I haven\u0026rsquo;t used it myself; needs verification\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003e\u003cstrong\u003eNeeds further investigation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRedCave: \u003ca href=\"https://www.redcave.com/\"\u003ehttps://www.redcave.com/\u003c/a\u003e\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eStill in beta; not sure which chain it\u0026rsquo;s on\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eIMO Blue Cat Digital: \u003ca href=\"https://www.lanmsz.cn/\"\u003ehttps://www.lanmsz.cn/\u003c/a\u003e\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eConsortium chain. Not sure about third-party integrations. They claim not to rely on any platform but use a consortium chain hhh. However, this seems to be combined with game IPs; let\u0026rsquo;s see how their world-building develops.\u003c/p\u003e","title":"Common NFT Websites"},{"content":"Since the content is indeed scarce, I searched the entire web, and there aren\u0026rsquo;t many people working on Deep Learning for Pixel Art either; there must really be no market for it, hhhh.\nFirst, let\u0026rsquo;s talk about the ones on GitHub labeled with pixel-art-generator or image-to-pixel. The former are mostly randomly generated, meaningless pixel blocks, while the latter mostly just apply OpenCV or some other image processing library resize, which is basically it. At best, they might add a filter, which is already impressive.\nI can only say that pixel art style is somewhat abstract yet requires very specific elements, which is likely something we still can\u0026rsquo;t fully handle right now.\nWhat is Pixel Art Regarding the definition of pixel art, I\u0026rsquo;ve heard many versions, but I feel this varies from person to person,since pixel art styles are evolving; pixel characters drawn years ago already differ significantly from today\u0026rsquo;s styles.\nFirst, there is a consensus: no anti-aliasing, meaning every pixel is distinct, with no blurry boundaries. After all, pixel art is an art form based on pixels, requiring pixel-level modifications.\nThe rest, I feel, depends on personal preference. Some like borders, others prefer no borders, half-borders, or discontinuous/non-borders.\nSecondly, regarding pixel art colors, there are actually no strict requirements now. Previously, computer limitations restricted us to 8-bit colors, but now you can use almost any color, though certainly not too many (this mainly applies to small-scale works).\nGalleries \u0026amp; Datasets Before we begin, let\u0026rsquo;s take a look at the galleries and datasets!\nGalleries eBoy: https://hello.eboy.com/pool/everything/1\nMany of the works above are excellent!\nPixelJoint: https://pixeljoint.com/\nThis is a gallery from a pixel art forum; the works are a mix of good and bad.\nOpenGameArt: https://opengameart.org/art-search?keys=pixelart\nFiltering for pixel art yields quite a few works.\nspriters-resource: https://www.spriters-resource.com/nes/\nContains some sprites and game screenshots.\nprobertson: https://probertson.tumblr.com/\nAll include creation details.\nDatasets sprites: https://paperswithcode.com/dataset/sprites https://spritedatabase.net/download\nSome datasets on GitHub\nhttps://github.com/AgaMiko/pixel_character_generator/blob/master/data.zip https://github.com/avidvid/OUP/tree/master/Downloads/Characters Excellent Projects Let\u0026rsquo;s first look at some excellent projects!\ninfo\nThis section covers two topics: `Pixelate`, which is pixelization, and `Depixelate`, which is de-pixelization. Pixelate info\nBelow are some papers I haven't read yet, for reference only! Only some of the work uses deep learning, so I won't distinguish between them here. Also, I won't distinguish between projects and papers; I'll list them together. Papers will be enclosed in book title marks. 《Automatic portrait image pixelization》 1\nThis is a 2021 paper. Actually, some information is lost; it only uses image processing, implemented in MATLAB. Unfortunately, the code doesn\u0026rsquo;t seem to be public, but it doesn\u0026rsquo;t look too complex.\n《Pixelated image abstraction》 2\nThe results look quite good. This is a 2012 paper, but it actually doesn\u0026rsquo;t use deep learning.\npixel_character_generator 3\nUses DCGAN, Conditional DCGAN, and DC AutoEncoder for character generation; I can only say the results are not ideal.\nMake Pixel Art in Seconds with Machine Learning 4\nThis uses CycleGAN. Actually, the results are okay. It mentions that training with cartoon images yields better results than real-world images, which makes sense since the cartoon domain is closer to pixel art. Here I\u0026rsquo;ll also paste cartoonset\npixel-me 5 [demo]\nThe results are indeed awesome, but it seems mainly targeted at faces; the effect on other domains is just average. Although there\u0026rsquo;s no paper or code, it likely involves background removal, generation with Pix2pix, and finally adding outlines.\n《Deep Unsupervised Pixelization》 6 7 8\nThis is actually work published at SIGGRAPH Asia 2018, np. It achieves pixelization via an unsupervised method; I haven\u0026rsquo;t studied the specifics yet.\neBoyGAN 9 [colab]\nThe author trained using StyleGAN with data from the aforementioned eBoy dataset. There\u0026rsquo;s a pre-trained model on Colab, but it seems it no longer runs.\nDepixelate There are actually many works on de-pixelization, but most are not done using deep learning, and there isn\u0026rsquo;t a very good compilation of them.\n《MMPX Style-Preserving Pixel Art Magnification》 10 11\nAlso provides a Web tool with implementations and comparisons of various algorithms 12 Here is my own implementation13, but to be honest, I feel the results are just average; it might not even be as good as xBR2X14\n《Geometric Total Variation for Image Vectorization, Zooming and Pixel Art Depixelizing》 15 demo 16\nAlso doesn\u0026rsquo;t use deep learning; it directly vectorizes the image, np. Maybe everyone feels that going from pixel art to normal graphics doesn\u0026rsquo;t require deep learning, hhh.\ninfo\nThe algorithms provided in these two papers are worth referencing. Other Works 《Towards Machine-Learning Assisted Asset Generation for Games: A Study on Pixel Art Sprite Sheets》 17 18\nThis work mainly uses Pix2pix to colorize pixel art, offering a good approach: first handle the light-dark relationships, then perform semantic segmentation on the characters, and finally generate characters in various colors (very suitable for making NFTs, hhhh). Unfortunately, the code is not open-sourced, and the dataset is not public.\nDrawing Tools There are countless tools for drawing pixel art. After all, this is a dimensionality reduction strike by other image editors. Besides traditional image editors that basically support pixel art, the most professional and widely used tool currently is aseprite.\naseprite: https://github.com/aseprite/aseprite\nAlthough this is open source, it requires self-compilation or purchase on Steam.\nLibreSprite: https://github.com/LibreSprite/LibreSprite\nThis is derived from the last GPLv2 commit of aseprite, has a release, and is also quite good.\nThere are a few other projects that are just too much; as I mentioned before, the basic implementation isn\u0026rsquo;t difficult—it\u0026rsquo;s just reinventing the wheel over and over.\nPixelorama pixel-art-react Conclusion With the current results based on GANs, one easily observable phenomenon is 太脏了: there\u0026rsquo;s simply too much fine-grained noise. This aligns with what I said at the beginning: although pixel art is abstract, every single pixel is very specific. There shouldn\u0026rsquo;t be transitions with only minor color differences. While subtle color variations can introduce color changes, if they\u0026rsquo;re too continuous, the result loses that authentic pixel art feel, and no one would really acknowledge it as pixel art. As for excellent projects like pixel-me, besides applying some GAN-specific techniques, they also performed image processing before and after the GAN stage. However, they haven\u0026rsquo;t open-sourced their work, which is genuinely disappointing, and they\u0026rsquo;ve even made a paid software version :(\nIf I have time later, I think I\u0026rsquo;ll try to do something similar.\n[pdf]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[pdf]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[code]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[url]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[blog]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[pdf]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[sup]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[code]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[code]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[page]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[pdf]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[url]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nimplementation\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nxBR2X\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[pdf]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[code]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[pdf]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n[blog]\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-deep-learning-for-pixel-art/","summary":"\u003cp\u003eSince the content is indeed scarce, I searched the entire web, and there aren\u0026rsquo;t many people working on Deep Learning for Pixel Art either; there must really be no market for it, hhhh.\u003c/p\u003e\n\u003cp\u003eFirst, let\u0026rsquo;s talk about the ones on GitHub labeled with \u003ccode\u003epixel-art-generator\u003c/code\u003e or \u003ccode\u003eimage-to-pixel\u003c/code\u003e. The former are mostly randomly generated, meaningless pixel blocks, while the latter mostly just apply \u003ccode\u003eOpenCV\u003c/code\u003e or some other image processing library \u003ccode\u003eresize\u003c/code\u003e, which is basically it. At best, they might add a \u003ccode\u003efilter\u003c/code\u003e, which is already impressive.\u003c/p\u003e","title":"Deep Learning for Pixel Art: Research, Datasets, and Tools"},{"content":"Lately, I\u0026rsquo;ve been completely obsessed with drawing. After picking up some half-baked skills in thick painting, I\u0026rsquo;ve decided to go back to pixel art.\nI created a pixel art piece, tweaked the colors a bit, and listed it on Bigverse (NFT China). However, the review process was painfully slow, taking about three days. Initially, I thought Bigverse was a public chain and that my work should be upgradable to Ethereum, with Bigverse just handling the on-chain process on my behalf. I didn\u0026rsquo;t realize that while Bigverse isn\u0026rsquo;t a private chain like those run by tech giants, it is actually a sidechain. This was my first time hearing about such a thing. Although transactions can be recorded and verified, being a sidechain means they cannot be verified on Ethereum, so chain-based uniqueness cannot be guaranteed (I heard Bigverse requires a manual review after searching for the image on both Baidu and Google first\u0026hellip;). A sidechain merely handles some public chain tasks, so there\u0026rsquo;s no need for gas fees. It turns out this works on the same principle as Polygon on OpenSea. Wow, if I had known this earlier, why would anyone bother with Bigverse? Wouldn\u0026rsquo;t Polygon be much better?\nOverall, Bigverse still feels like it\u0026rsquo;s in development. The frontend interface is just terrible; I guess all the work is focused on the backend on-chain process. Also, the fuel fee is a bit high, yet they call it a \u0026lsquo;paid service\u0026rsquo; without telling us whose wallet it goes into. Furthermore, the platform deducts a transaction fee for every trade? Confusing? This already goes against the original intent of NFTs—charging fees just to mint and trade on the chain? Additionally, anonymity is unrealistic and likely non-existent in China, so the anonymity that blockchain promises is lost here.\nThat said, Bigverse is indeed one of the few available solutions in China where virtual currencies are effectively non-existent. If it\u0026rsquo;s a sidechain, so be it; at least it\u0026rsquo;s far better than private chains like Alipay\u0026rsquo;s that exist just to scam people? There are a few viable development paths for the future. One approach, which Bigverse seems to be pursuing now, is signing famous artists to mint their works. Another idea is to combine NFTs with physical goods; given the recent popularity of 3D-style NFTs, creating figurines via 3D printing seems feasible. Additionally, collaborations with social media platforms and game developers could work. Currently, the biggest use case for NFTs is probably profile pictures, hhhh. Social media could directly support profile pictures, while game developers could allow users to DIY game skins or character avatars. If NFTs become fully legal in China, this group of pioneers at Bigverse will definitely be the first to reap the benefits.\nFinally, if you register through this link, you can get 3 free fuel credits, meaning you can mint 3 works for free.\n","permalink":"https://blog.bj-yan.top/en/p/misc-bigverse-first-experience/","summary":"\u003cp\u003eLately, I\u0026rsquo;ve been completely obsessed with drawing. After picking up some half-baked skills in thick painting, I\u0026rsquo;ve decided to go back to pixel art.\u003c/p\u003e\n\u003cp\u003eI created a pixel art piece, tweaked the colors a bit, and listed it on \u003ca href=\"https://www.nftcn.com.cn/pc/#/mall/mallDetail?tid=91619467872643967426164977164512\"\u003eBigverse (NFT China)\u003c/a\u003e. However, the review process was painfully slow, taking about three days. Initially, I thought Bigverse was a public chain and that my work should be upgradable to Ethereum, with Bigverse just handling the on-chain process on my behalf. I didn\u0026rsquo;t realize that while Bigverse isn\u0026rsquo;t a private chain like those run by tech giants, it is actually a sidechain. This was my first time hearing about such a thing. Although transactions can be recorded and verified, being a sidechain means they cannot be verified on Ethereum, so chain-based uniqueness cannot be guaranteed (I heard Bigverse requires a manual review after searching for the image on both Baidu and Google first\u0026hellip;). A sidechain merely handles some public chain tasks, so there\u0026rsquo;s no need for gas fees. It turns out this works on the same principle as Polygon on OpenSea. Wow, if I had known this earlier, why would anyone bother with Bigverse? Wouldn\u0026rsquo;t Polygon be much better?\u003c/p\u003e","title":"First Experience with Bigverse"},{"content":"The idea started when I saw a comment request on Course Evaluation Website . Having previously used giscus, I briefly checked the MkDocs documentation and conveniently submitted a PR to add the comment system. I then had the idea of building such a repository for Hainan University, adhering to the principle of open sharing on the internet to share course materials.\nThus, this repository was created.\nA minor modification to the GitHub Action workflow allowed it to go live directly. I shared the news in several groups (fortunately, I didn\u0026rsquo;t get banned).\nAs of 2021-10-28 00:00:00, on the first day of the website launch, it received 474 page views, gained 255 new users, and collected 8 stars and 4 forks!!\nThanks to the students who promoted the website on the first day; the website\u0026rsquo;s growth relies on everyone\u0026rsquo;s continued attention!\nI hope this project can continue running. I expect heavy traffic during course registration and final exam periods, but daily maintenance is the foundation for the website\u0026rsquo;s sustainable development.\n","permalink":"https://blog.bj-yan.top/en/p/blog-hainanu-course-comments/","summary":"\u003cp\u003eThe idea started when I saw a comment request on \u003ca href=\"https://github.com/conanhujinming/comments-for-awesome-courses\"\u003eCourse Evaluation Website \u003c/a\u003e. Having previously used giscus, I briefly checked the MkDocs documentation and conveniently submitted a PR to add the comment system. I then had the idea of building such a repository for Hainan University, adhering to the principle of open sharing on the internet to share course materials.\u003c/p\u003e\n\u003cp\u003eThus, this \u003ca href=\"https://github.com/yyyanbj/hainanu-course-comments\"\u003erepository \u003c/a\u003e was created.\u003c/p\u003e\n\u003cp\u003eA minor modification to the GitHub Action workflow allowed it to go live directly. I shared the news in several groups (fortunately, I didn\u0026rsquo;t get banned).\u003c/p\u003e","title":"Hainan University Course Guide Sharing Initiative"},{"content":"Preface Today I watched 3b1b1\u0026rsquo;s \u0026lsquo;Essence of Linear Algebra]2\u0026rsquo;, which gave me many new insights and clarified many things I hadn\u0026rsquo;t fully understood before (or rather, formulas I had blindly memorized just to pass exams). I couldn\u0026rsquo;t wait to write this article. The geometric intuitions used here do not involve formal mathematical proofs.\nKnowledge Matrix Multiplication First, let me record some key points mentioned in the video. There\u0026#39;s nothing much to say about vectors themselves; instead, vector transformations lead us to matrix multiplication. Vector transformations are essentially coordinate system transformations, or what we call linear transforms. Consider a vector $\\overrightarrow{v}=[x\\hat{i}, y\\hat{j}]^{T}$. Applying a transformation $A=\\begin{bmatrix}a \u0026amp; c\\\\ b \u0026amp; d\\end{bmatrix}$ to it: $\\hat{i},\\hat{j}$ represents the basis vectors. The transformation applied to $\\hat{i}$ is $\\begin{bmatrix}a\\\\b\\end{bmatrix}$, and the transformation applied to $\\hat{j}$ is $\\begin{bmatrix}c\\\\d\\end{bmatrix}$. The coordinates after the transformation are $\\begin{bmatrix}a \u0026amp; c\\\\ b \u0026amp; d\\end{bmatrix}\\begin{bmatrix}x\\\\ y\\end{bmatrix}=x\\begin{bmatrix}a\\\\ b\\end{bmatrix}+y\\begin{bmatrix}c\\\\ d\\end{bmatrix}=\\begin{bmatrix}ax+cy\\\\bx+dy\\end{bmatrix}$ Finally, I don\u0026rsquo;t have to memorize the row-column order for matrix multiplication!\nThis extension to higher dimensions works the same way. Below, let\u0026rsquo;s consider multiple transformations. Applying two transformations $A,B$ (which is equivalent to $AB\\overrightarrow{v}$) is essentially a composite function $A(B\\overrightarrow{v})$: first performing a $B$ transformation, followed by a $A$ transformation. Why can\u0026rsquo;t we swap this order? A more intuitive explanation is that after the first $B$ transformation, the coordinate system has already changed. Although $A$ transforms coordinates in the same way, its effect on the already-transformed $B\\overrightarrow{v}$ is different from its effect if applied directly to $\\overrightarrow{v}$. This explains why matrix multiplication is not commutative.\nWhat about associativity? Notice that both $(AB)C\\overrightarrow{v}$ and $A(BC)\\overrightarrow{v}$ are actually associated from right to left. The result of multiplying the first two matrices is also combined in order, so the associative law does hold (even if the intuitive feeling might seem strange here, this is indeed the correct transformation sequence).\nIt is my experience that proofs involving matrices can be shortened by 50% if one throws the matrices out. - Emil Artin\nDeterminant Previously, my learning of determinants was biased; I only used them to solve systems of linear equations and never considered their geometric meaning. Here, the geometric meaning of a determinant is the factor by which the unit area changes under a \u0026rsquo;transformation\u0026rsquo;—essentially, the \u0026lsquo;area\u0026rsquo; of the unit square after the transformation. This actually made me pause for a moment. If we accept the above statement, then the area after the matrix transformation represented by $[1,1]^{T}$ should be the determinant. Following this logic, $\\begin{bmatrix}a \u0026amp; c\\\\b \u0026amp; d\\end{bmatrix}\\begin{bmatrix}1\\\\1\\end{bmatrix}=\\begin{bmatrix}a+c\\\\b+d\\end{bmatrix}$, the area of the identity matrix should be $(a+c)(b+d)=ab+ad+bc+cd$. However, we all know the value of a 2x2 determinant should be $ad-bc$. Why the discrepancy? It\u0026#39;s simply because \u0026#39;the coordinate system changed.\u0026#39; $ad-bc$ uses the original coordinate system, while $(a+c)(b+d)$ uses the transformed one. So, can a determinant be negative? Yes, this corresponds to a flip in the coordinate system. For example, rotating the plane coordinate system $x$ to $y$ by 90° counterclockwise, or changing the $xyz$ coordinate system from a right-handed system to a left-handed system, will both result in a negative determinant. For higher dimensions, fixing all other dimensions and transforming two of them will make the determinant negative; transforming another dimension makes it positive again, then negative again, and so on. However, this seems impossible to prove geometrically, as the human brain struggles to visualize high-dimensional spaces.\nWhat if the determinant of a transformation becomes 0? The unit area becomes 0, meaning the transformation has been \u0026lsquo;dimensionally reduced\u0026rsquo; into a line. This implies that the column vectors of the transformation are not linearly independent.\nAt the same time, 3b1b left a thought-provoking question $\\det(M_1M_2)?=\\det(M_1)\\det(M_2)$. The result is, of course, equal. 3b1b did not provide a geometric explanation, but here is my understanding: it is essentially multiplication. Looking at the right side, following the right-to-left association order, we first scale the original unit matrix by $\\det(M_2)$, then apply the $M_1$ transformation to scale it to $\\det(M_1)$, obtaining the final unit area. This is equivalent to performing two transformations sequentially, reorganizing the coordinates after each transformation before proceeding to the next. On the left side of the equation, it\u0026rsquo;s as if both transformations were performed at once, directly scaling to $\\det(M_1)\\det(M_2)$, thus yielding $\\det(M_1M_2)=\\det(M_1)\\det(M_2)$.\nInverse Matrix \u0026amp; Rank The inverse matrix, or inverse \u0026ldquo;transformation,\u0026rdquo; brings us to a point where all matrices can essentially be referred to as \u0026ldquo;transformations,\u0026rdquo; because what we are dealing with here are coordinate system transformations, i.e., linear transforms. One point I previously forgot to mention is that the result of a coordinate system transformation caused by a linear transform is always linear, meaning the coordinate axes remain \u0026ldquo;straight,\u0026rdquo; must be \u0026ldquo;parallel,\u0026rdquo; and are also \u0026ldquo;equidistant.\u0026rdquo; If these conditions are not met, a straight line in the original coordinate system would no longer be a straight line, and thus it would no longer be a linear transformation.\nAs an inverse transformation, it is written as $A^{-1}$. Applying a transformation and then transforming it back is essentially $A^{-1}A$ (which might explain why we always write $A^{-1}$ on the left? Because the combination always starts from the right). Solving an equation $Ax=v$ is essentially finding a transformation that maps $v$ to $x$. Note that it is not transforming $x$ into $v$, because $x$ is the unknown variable. This is essentially $A^{-1}Ax=A^{-1}v$, and the solution is $x=A^{-1}v$.\nIf $\\det(A)=0$, then the transformation that has been subjected to \u0026ldquo;dimensional reduction\u0026rdquo; can never be restored to its pre-transformation state. This is equivalent to saying you can never find a plane from two parallel line vectors, meaning no inverse exists, yet there are infinitely many solutions.\nAfter transformation, the minimum dimension it compresses to is the rank. This is derived from the transformation of each column vector, resulting in the column space.\nRegardless of any transformation, the vector $\\begin{bmatrix}0\\\\0\\end{bmatrix}$ must remain at the origin. For a full-rank matrix, only the $[0,0]^{T}$ vector stays at the origin. For a non-full-rank matrix, a series of vectors will be compressed until they reach the $[0,0]^{T}$ vector; these constitute the \u0026#34;null space\u0026#34; or \u0026#34;kernel.\u0026#34; Non-square Matrices Considering the \u0026ldquo;column space,\u0026rdquo; left-multiplying by a non-square matrix is essentially performing a projection. If you left-multiply by a square matrix of size $3\\times 2$, it is essentially transforming and projecting a 2D plane into 3D space.\nThe dot product itself has geometric meaning; it can also be viewed as the projection described above, projecting onto a one-dimensional space vector.\nMore thing Let\u0026rsquo;s think about other expressions within the matrix.\n$(AB)^{-1}=B^{-1}A^{-1}$ .\nFrom right to left, first combine the $B$ transformation, then combine the $A$ transformation. To reverse this process, considering that the coordinate system must remain stationary, we must first apply the inverse of the $A$ transformation, and then apply the inverse of the $B$ transformation.\n3b1b\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nhttps://space.bilibili.com/88461692/channel/detail?cid=9450\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/blog-linear-algerbra-ng/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eToday I watched 3b1b\u003csup id=\"fnref:1\"\u003e\u003ca href=\"#fn:1\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e1\u003c/a\u003e\u003c/sup\u003e\u0026rsquo;s \u0026lsquo;Essence of Linear Algebra]\u003csup id=\"fnref:2\"\u003e\u003ca href=\"#fn:2\" class=\"footnote-ref\" role=\"doc-noteref\"\u003e2\u003c/a\u003e\u003c/sup\u003e\u0026rsquo;, which gave me many new insights and clarified many things I hadn\u0026rsquo;t fully understood before (or rather, formulas I had blindly memorized just to pass exams). I couldn\u0026rsquo;t wait to write this article. The geometric intuitions used here do not involve formal mathematical proofs.\u003c/p\u003e\n\u003ch2 id=\"knowledge\"\u003eKnowledge\u003c/h2\u003e\n\u003ch3 id=\"matrix-multiplication\"\u003eMatrix Multiplication\u003c/h3\u003e\n\u003cdiv class=\"has-mathjax\"\u003e\n\n\nFirst, let me record some key points mentioned in the video. There\u0026#39;s nothing much to say about vectors themselves; instead, vector transformations lead us to matrix multiplication. Vector transformations are essentially coordinate system transformations, or what we call linear transforms. Consider a vector $\\overrightarrow{v}=[x\\hat{i}, y\\hat{j}]^{T}$. Applying a transformation $A=\\begin{bmatrix}a \u0026amp; c\\\\ b \u0026amp; d\\end{bmatrix}$ to it: $\\hat{i},\\hat{j}$ represents the basis vectors. The transformation applied to $\\hat{i}$ is $\\begin{bmatrix}a\\\\b\\end{bmatrix}$, and the transformation applied to $\\hat{j}$ is $\\begin{bmatrix}c\\\\d\\end{bmatrix}$.\nThe coordinates after the transformation are $\\begin{bmatrix}a \u0026amp; c\\\\ b \u0026amp; d\\end{bmatrix}\\begin{bmatrix}x\\\\ y\\end{bmatrix}=x\\begin{bmatrix}a\\\\ b\\end{bmatrix}+y\\begin{bmatrix}c\\\\ d\\end{bmatrix}=\\begin{bmatrix}ax+cy\\\\bx+dy\\end{bmatrix}$\n\n\n\u003c/div\u003e\n\u003cblockquote\u003e\n\u003cp\u003eFinally, I don\u0026rsquo;t have to memorize the row-column order for matrix multiplication!\u003c/p\u003e","title":"Investigating Things to Extend Knowledge - Linear Algebra"},{"content":" The essence of the world is mathematics.\nSince \u0026lsquo;Yuan Gui\u0026rsquo; landed in Hainan, it has weakened considerably; the outside world is no longer experiencing crazy, fierce gales. It should have passed its most intense phase by now.\nA few days ago, I went to the library and found that, as expected, they had purchased many new books. I wonder if it was because of my suggestion 233 (I seem to have filled out this feedback when the library previously asked for input).\nRecently, I\u0026rsquo;ve also read some books on optimization and happened to re-review SGD. Previously, I just skimmed the books and it was simple to get through, but when I tried to derive it by hand, I realized I still didn\u0026rsquo;t fully understand it. The reason is that I never systematically studied matrix calculus; in many cases, I couldn\u0026rsquo;t grasp the derivation process and merely applied formulas by rote. Therefore, I recently planned to read Matrix Cookbook1!\nMatrix Cookbook\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/misc-20211013/","summary":"\u003cblockquote\u003e\n\u003cp\u003eThe essence of the world is mathematics.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eSince \u0026lsquo;Yuan Gui\u0026rsquo; landed in Hainan, it has weakened considerably; the outside world is no longer experiencing crazy, fierce gales. It should have passed its most intense phase by now.\u003c/p\u003e\n\u003cp\u003eA few days ago, I went to the library and found that, as expected, they had purchased many new books. I wonder if it was because of my suggestion 233 (I seem to have filled out this feedback when the library previously asked for input).\u003c/p\u003e","title":"The Essence of the World is Mathematics"},{"content":"Indeed, this was the first time since I turned 21 that I called 120 and rode in an ambulance. When my elders were sick before, I neither rode in an ambulance nor made the 120 call, and I might not have even been the first to know.\nWhat made this time different was that the situation involved my roommate. After calling 120, the first question they asked was the address, then the patient\u0026rsquo;s condition, and surprisingly, they even asked about pandemic-related details\u0026hellip; - -\nMy roommate likely suffered some lung injury, leading to breathing difficulties. As a result of panic, the condition worsened; after breathing heavily with his mouth wide open, he began to experience numbness, starting from his face, spreading to his extremities, and then to his limbs. At that point, he was drenched in cold sweat, struggling to breathe, and said, \u0026lsquo;Xiao Bing, I feel like I\u0026rsquo;m done for.\u0026rsquo; Then he started gasping. He had no medical history; it was purely due to the physical fitness test. I guess he couldn\u0026rsquo;t catch his breath during the 1000-meter run or perhaps had a ruptured alveolus or something similar. Anyway, he kept putting it off until the 6th, when he finally felt unwell, returned to the dorm, and luckily, I, the idle person, was there (233).\nWhile waiting for the ambulance, I guided them on how to get there. Upon arrival, it turned out there was no major issue; it was just the heavy breathing with his mouth open. If he had taken deep breaths through his nose instead, it would have been much better. On the way to the hospital, we happened to encounter a minor traffic accident on campus where someone was injured, so we provided some first aid. This was also my first time riding in an ambulance\u0026hellip; - -, and in the end, the bill was 115 yuan.\nAs long as there\u0026rsquo;s no major issue, that\u0026rsquo;s good. This incident teaches us to seek medical attention early if we\u0026rsquo;re sick or feel unwell, to avoid delaying until serious problems arise\u0026hellip; - -.\n","permalink":"https://blog.bj-yan.top/en/p/misc-20211006/","summary":"\u003cp\u003eIndeed, this was the first time since I turned 21 that I called 120 and rode in an ambulance. When my elders were sick before, I neither rode in an ambulance nor made the 120 call, and I might not have even been the first to know.\u003c/p\u003e\n\u003cp\u003eWhat made this time different was that the situation involved my roommate. After calling 120, the first question they asked was the address, then the patient\u0026rsquo;s condition, and surprisingly, they even asked about pandemic-related details\u0026hellip; - -\u003c/p\u003e","title":"My First Time Calling 120 \u0026 My First Ride in an Ambulance"},{"content":"Preface September 30, 2021. Finally, the weight on my heart has been lifted.\nAlthough I had long known the outcome was a done deal, only when I clicked \u0026lsquo;Accept\u0026rsquo; in the system did I feel that everything had finally settled.\nThis is already my fourth year of university life. From that initially withdrawn, introverted, and taciturn high school student, I have gradually grown into who I am today.\nOnly now, feeling that my efforts have finally yielded corresponding rewards, do I feel that everything was worth it.\nBefore University I am a student from Shandong. I can say that my academic performance in regular subjects was never particularly good, far behind the top students. In middle school, relying on a bit of \u0026lsquo;cleverness,\u0026rsquo; I stayed around the top 3 in my class, though it wasn\u0026rsquo;t very stable later on. I successfully got into Shengli No. 1 Middle School for the high school entrance exam, which was considered a great success at the time 233 (by the way, the nickname \u0026lsquo;Xiao Bing\u0026rsquo; was given to me back then in middle school; I also had a friend named \u0026lsquo;Nandao\u0026rsquo;).\nHowever, there is always someone better, and mountains beyond mountains. The advantages I had in middle school were completely gone. I must admit, there are indeed significant differences in educational standards across regions, much like the difference between \u0026lsquo;giving a fish\u0026rsquo; and \u0026rsquo;teaching how to fish.\u0026rsquo;\nIn high school, the luckiest thing that happened to me was participating in the informatics competition. Coincidentally, after learning about informatics competitions for just two days in middle school, I dared to apply for the independent recruitment exam at Shengli No. 1. At the time, fueled by sheer determination, I read through Tan Haoqiang\u0026rsquo;s \u0026lsquo;C Programming Design\u0026rsquo; and felt that finishing it was a success in itself. Of course, independent recruitment or competitions didn\u0026rsquo;t test these basic syntaxes; algorithms were completely unknown to me back then, or rather, I had never even heard of them. The result was predictable, and to this day, I still feel I was truly bold.\nBut this turned out to be a stroke of fate. I successfully left my name in the informatics group, which became the opportunity for me to join. Shengli No. 1\u0026rsquo;s tradition was to ask applicants to fill in their areas of interest when submitting materials (I don\u0026rsquo;t quite remember the specific method). If your ranking was high enough, they would call to ask if you were willing to participate in competition training, which would take place during the summer before high school started. My high school entrance exam scores weren\u0026rsquo;t particularly good, but the name I had left earlier played a role. My mom didn\u0026rsquo;t receive a text message notification (since I didn\u0026rsquo;t have a phone at the time, and the contact information provided was all for my parents), but another parent who knew my mom saw my name on the list in a group chat and told her (I\u0026rsquo;m ashamed to say, I can\u0026rsquo;t remember who that parent was). My mom immediately decided to take me to the training. To this day, I still don\u0026rsquo;t know why she was so insistent about making me go (at least looking back, my mom always cared deeply about my opinions). I packed my quilt and went to report the next day. At that time, I had just finished the high school entrance exam and was preparing to relax and slack off, so I was very reluctant. Perhaps it was because I was used to obeying, but I went anyway. When my mom arrived at the school, she spoke with the teacher in charge of the competition (Old Li), saying she hadn\u0026rsquo;t received a text message but my name was on the list. I didn\u0026rsquo;t hear the details of their conversation; I was just standing nearby, hoping the teacher would say I could go back. But after a while, I was told to go upstairs. That was the building to the right of the entrance of the old No. 1 High School. I don\u0026rsquo;t remember its name now, and it no longer exists (the old No. 1 High School was demolished).\nThat was another world. During the summer training, I truly came into contact with algorithms and competitions. I learned that so many students from the middle school department of No. 1 High School had already participated in competitions, and I discovered there was something called the Olympiad in Informatics and a competition called NOI(P). Since my home wasn\u0026rsquo;t in Dongying, I naturally had to board, and I built revolutionary friendships with other students who were also boarding.\nPerhaps seeing my small \u0026rsquo;talent,\u0026rsquo; I passed the first round of screening at the end of the summer, successfully skipping the military training (continuing with the training, though we had also had military training in middle school). I missed a great opportunity to bond with my classmates (although after that, I felt like I had become a \u0026lsquo;Compiler Principles person\u0026rsquo;).\nHowever, my path in competitions was not smooth; there are even some parts I don\u0026rsquo;t want to recall. I\u0026rsquo;ll briefly note them here, and if I remember anything later, I\u0026rsquo;ll come back to add more.\nFirst year: Attended training on regular weekends, attended normal classes otherwise. Result: Provincial Second Prize in NOIP.\nSecond year: Almost all subjects except core ones were dedicated to competitions. Result: Provincial First Prize in NOIP, Silver Medal in APIO, Bronze Medal in CTSC. Failed in the first round of provincial selection (if I remember correctly, I lost nearly 200 points due to two small mistakes; a slight modification afterward would have passed). Failed to turn things around in the second round (my mindset was truly broken during the first and second rounds, and I had no desire to solve problems). Retired from competitions.\nThird year: Retired from competitions, focused on catching up on regular subjects, took the college entrance exam (Gaokao), and applied for independent recruitment. Shandong University: -20 points, Tianjin University: -60 points. Gaokao score: 582, provincial rank approximately 30,000. Missed the mark. Gained admission to Hainan University with my raw score.\nTo say I wasn\u0026rsquo;t unwilling would be a lie. My competition results brought me no advantages; instead, they seemed to become a burden.\nLater, many people asked me if I regretted participating in the competitions. If I hadn\u0026rsquo;t, and had devoted myself entirely to regular subjects, my scores would certainly have been higher.\nHowever, my answer has never changed: \u0026lsquo;No regrets. I love it.\u0026rsquo;\nWithout the competitions, I don\u0026rsquo;t know how I would have spent my three years of high school. It would have been just reading and studying, reading and studying.\nThis is absolutely the most precious wealth of my life, perhaps without equal. Even considering the outcome, the answer remains the same. As someone who grew up in a small town, I was exposed to a much broader world, met impressive people, learned about their efforts, and recognized the gaps in talent and insight. Sometimes, a more comprehensive understanding of oneself is more important than anything else. At the same time, I felt the \u0026lsquo;gap\u0026rsquo;—not just in talent, but in environment and cognition.\nAlong this journey, whether the people I met were good or bad, I am grateful for them all.\nHowever, regarding my hometown, Shandong, although I shouldn\u0026rsquo;t feel this way, I truly don\u0026rsquo;t like it. In the fierce struggle of the Gaokao, my desperate desire to get out was largely due to this reason. I wanted to go south. In my college application, I probably only listed Shandong University as a reach to see if I could make it, and Qingdao University as a safety net. Apart from those, all my choices were schools in the south. The final result was also a 211 university: Hainan University.\nHainan University First Encounter Hainan University, a bottom-tier 211. Why it managed to get the 211 designation, those who know probably know the reasons.\nWhen I first arrived at Hainan University, I saw dilapidated dormitory buildings and shabby rooms. The door handles and bed frames were covered in rust (although I later learned that our dormitory building is actually called the \u0026lsquo;Prince Building\u0026rsquo; because it can be considered the best male dorm in Hainan University: 4-person rooms with private bathrooms. I\u0026rsquo;m very grateful for our counselor\u0026rsquo;s luck at the time!). Later, after comparing with dorms at other schools, I realized ours were actually quite comfortable,after all, the space was large, with one bed and desk per person). But I never considered repeating a year; I didn\u0026rsquo;t want to go back and study those boring cultural subjects again. I just adopted a \u0026lsquo;go with the flow\u0026rsquo; mindset and came here.\nBut after staying here for a while, I actually grew to like it. I probably fell in love with the \u0026lsquo;coziness\u0026rsquo; here. Now, if anyone says Hainan University is bad, I will definitely be the first to jump out and oppose them!\nUniversity is where friends from all over gather. I\u0026rsquo;m fortunate that my three roommates lived with me in harmony; our dormitory relationship never had any issues. Everyone was quite generous. I love you, hhh.\nTo be honest, when I first arrived at Hainu, I indeed held a very dissatisfied attitude, even feeling that this place didn\u0026rsquo;t deserve me, which led me to say many negative things. However, one thing that can be pointed out is that Shandong students coming to Hainan University actually have quite high scores compared to students from other provinces. Especially when Shandong uses the National Paper I, while other provinces use National Papers II or III, yet the cutoff scores are similar.\nAs a freshman, with no interview experience whatsoever, and even quite introverted, I still signed up for many student organizations, including the University Youth Volunteer Association, University Youth League, College Youth League, College Youth Volunteer Association, and many others I\u0026rsquo;ve forgotten. But I forgot them back then, not now. This also led to many interesting events.\nAlthough I knew nothing and the future was uncertain, and although I was actually the first college student in my family with no experience to guide me from home (some advice was even wrong, just hearsay and simple herd mentality), the one thing I knew for sure was that university was the start of my new life. I also genuinely wanted to change myself from the bottom of my heart. It was this mindset of daring to participate and daring to try that brought me some opportunities.\nThe outcome for an introverted person with little experience was predictable. Sometimes when attending interviews, I would only receive the location and time information. Sometimes I even mixed up which organization was holding the interview\u0026hellip; I hadn\u0026rsquo;t prepared my self-introduction much. If I were to interview myself back then today, I probably wouldn\u0026rsquo;t leave any impression, the kind of person who leaves no impression at all, hhh. Oh right, sometimes, truly, I didn\u0026rsquo;t know, myself, which, department, I, had, actually, filled in, hahahahahahaha. But luckily, the ministers of my College Youth Volunteer Association didn\u0026rsquo;t care how much you bragged; they valued your personality and character more. So I successfully joined the College Youth Volunteer Association. When I learned there was also a University Youth Volunteer Association, I instantly felt my level dropped significantly. But later, after seeing the workload of the University Youth Volunteer Association and watching my roommate who was in that org being so busy, I realized the student organization bonus points were actually the same. I even felt a bit relieved, hhhh.\nThrough participating in volunteer activities, writing articles, and taking photos over the year, I came to know and master some practical skills. Perhaps because I had a quite pleasant year, I finally chose to stay on in the position under the minister\u0026rsquo;s persuasion (perhaps they thought I was reliable, as others in our department also took the retention assessment, but I was ultimately selected).\nOne could say I forced myself not to be idle during my freshman year; I would find things to do whenever I had nothing to do. Another turning point also happened during this year.\nTowards the end of the first semester, we got a new counselor. The new counselor selected students interested in research. Since the counselor was at the State Key Laboratory, I naturally joined in. Under the recommendation of our class monitor, I became some kind of research leader (I completely can\u0026rsquo;t remember what title). As for why I was recommended? Probably because in the first year, the college had an \u0026lsquo;IT Culture Festival\u0026rsquo; with a programming competition. Relying on some algorithm knowledge I learned in high school, I easily achieved a perfect score (AK) and soundly defeated other contestants from the sophomore and junior years, hhhh (But to be honest, the difficulty of that competition\u0026rsquo;s problems was probably around NOIP Day 1 Problem 1 or Day 2 Problem 1 level. Most were just programming ability tests; if you could understand the problem, you could solve it. And there was no automated judging system; judging was done manually. You\u0026rsquo;d say you finished, someone would bring over test cases, and if the answers were correct, they\u0026rsquo;d give you points and record your time). From then on, everyone thought I was quite capable, especially in programming (who says competitions are useless, x).\nAdditionally, to add a note: when school first started, I also considered joining the ACM team. But after asking senior students, I didn\u0026rsquo;t even know what ACM was, so I gave up. Hainan University indeed didn\u0026rsquo;t have an ACM program. One reason is inconvenient transportation; traveling between cities requires flights since there are no high-speed trains, making the cost very high. Another reason I guess is that students\u0026rsquo; starting point was quite late, and there was no one to organize things, no one to pull you in for a summer training camp before the semester started. However, nowadays, the Robotics and Artificial Intelligence Association has an ACM group. That\u0026rsquo;s a story for another time.\nAnyway, I don\u0026rsquo;t even know what I was doing when I first entered the lab. Usually, I just took my shift, which basically meant dragging my laptop over there and studying by myself. But at that time, there was an upperclassman working on network security who organized some training sessions. I picked up a bit of ESP82661, 2, 3 and Arduino4 stuff, which was quite fun 233. Later, they organized the first HDCTF5, and that was my first time encountering CTF. Relying on the number theory knowledge I had from high school competitions 6, 7, I could understand and work through RSA and similar problems. Also, since I had previously done Nazo, some of the web challenges were easy to solve. At the time, I even deliberately studied HTML8 and JavaScript9 content, though I was really not good at anything else. In the end, I won third prize? I forgot, but I remember the award was presented by our homeroom teacher, and the prize was a wristband. After that, I tried participating in CTFs because I found them extremely interesting, mainly just for fun. Later, an upperclassman invited me to become the Vice President of the Cyberspace Security Association, and I happily agreed (although I mostly just ate a meal and slacked off, pushing tasks whenever possible, which is truly embarrassing). Later, due to some matters, I stopped participating in CTFs, which is the turning point I mentioned.\nThinking about it, I\u0026rsquo;d better not mention real names. At the lab, I met classmate HY, who then recommended me to BY. I learned that BY had already graduated from Hainan University (in fact, during military training, a video of him being named \u0026lsquo;Person of the Year\u0026rsquo; was even played at Siyuan Hall). He also had his RA-Team. This counts as another coincidence. If I hadn\u0026rsquo;t participated in the IT Culture Festival, if I hadn\u0026rsquo;t joined the counselor\u0026rsquo;s group, my class monitor wouldn\u0026rsquo;t have recommended me to be the xx person in charge, I wouldn\u0026rsquo;t have met HY, and thus I wouldn\u0026rsquo;t have known BY.\nAt that time, BY asked me a few simple questions, like what I could do, to get a basic understanding. After that, I somewhat confusedly joined my first innovation and entrepreneurship competition (the Hainan Province Artificial Intelligence Creativity Competition). My role was to present the PPT; I didn\u0026rsquo;t need to defend it, just memorize the script and present. I also met upperclassmen XM and LZ who were participating in the same competition. The result was quite good, roughly a first prize.\nThis story marks the official beginning.\nBusy It was from then on that I learned about innovation and entrepreneurship competitions and also understood the concept of recommendation for postgraduate studies without entrance exams. Although I hadn\u0026rsquo;t deliberately tried to boost my GPA, due to a lifelong stereotype, I still hoped for good grades, so I studied extremely hard during final exam review. My freshman year GPA was the peak of my entire university career, around 3.73. By June and July, most competitions had ended, and to participate again, I\u0026rsquo;d have to wait another year.\nDuring the summer break, two things happened. One was that BY arranged for me to learn machine learning, so I followed Professor Hongyi Li\u0026rsquo;s machine learning videos. I was probably watching the 2016 version at the time, and I had covered most of the material. I submitted my study notes to BY, and the two blog posts about machine learning notes 10, 11 were written then. In high school, I had actually been exposed to some Python, but programming languages have never been a barrier for me. At the time, BY also assigned several small projects, one of which was Raspberry Pi for MNIST12, which was my first encounter with the Raspberry Pi. My very first small project was Raspberry Pi for Fire Recognition13. Back then, I used Keras\u0026rsquo;s CNN; I didn\u0026rsquo;t understand things like VGG, so I just went straight with CNN. I also didn\u0026rsquo;t know about object detection; I just needed to determine whether there was fire or not. I once wanted to publish a paper based on this, focusing on sensor and vision fusion detection, but it was heartbreaking due to the lack of data. At the time, I didn\u0026rsquo;t even know what multimodal meant; it was just a basic idea. Later, this actually participated in several small competitions and served as the opening project for the machine learning group.\nThe other thing during the summer break, although my involvement was minimal, had a huge impact later on. It was the establishment of the Hainan University Robotics and Artificial Intelligence Association.\nWhen sophomore year began, everything got back on track. The busiest time at the start of the semester was welcoming new students and recruiting members (don\u0026rsquo;t forget I stayed on as the Minister of the News and Publicity Department of the College Youth Volunteer Association; of course, as Vice President of the Cyberspace Security Association, I also had to recruit; as the founder of the Robotics and AI Association, I also had to recruit; as the leader of the machine learning group, I also had to recruit\u0026hellip;). The association recruitment also brought in the first batch of staff and ministers. I also saw some interesting juniors there. Since there were basically no competitions in the previous semester except for the initial period, everything remained smooth and steady.\nNot long after the semester started, CJ, XM, and I went to Xi\u0026rsquo;an to participate in the Third Silk Road Robot Creativity Competition. Almost everyone who went won an award; except for the last few who got second prizes, the rest won first prizes and special prizes. We also saw some very impressive projects. In the end, we won a first prize and earned a fully funded trip to Xi\u0026rsquo;an.\nAlso, near the end of the first semester of sophomore year, BY told me about a new field called FL, noting that many problems remained unsolved, with incentive mechanisms being a key point. This later became the lead-in.\nNext came the unexpected pandemic.\nThis unprecedented pandemic disrupted everyone\u0026rsquo;s rhythm, and all cities came to a standstill. Ironically, it\u0026rsquo;s both lucky and laughable that Dongying never had a single case. Perhaps it\u0026rsquo;s because Dongying sits at the end of all transportation routes; the flow of people was never large to begin with, and control measures were relatively effective.\nThe pandemic was agonizing, yet it was also an opportunity. After all, the work for both of my papers was related to the pandemic, setting the stage for the subsequent competitions.\nHainan University was relatively bold; as soon as the pandemic was effectively controlled in May, they reopened. I originally thought the competition would be postponed due to the pandemic, but it wasn\u0026rsquo;t. With hasty notifications and rushed preparations, I inexplicably accepted the unmanned vehicle project, while handing over the FL-based pneumonia detection project I was working on during the pandemic to YZ. The result? Both projects performed poorly, while the other projects from the association won so many awards they were hard to hold. Reflecting later, the reasons were clear: I simply lacked experience and did a poor job, especially not understanding the judging rules or how to package the work properly (in fact, since then, I\u0026rsquo;ve developed some aversion to such competitions; they weren\u0026rsquo;t truly what interested me).\nHowever, as the saying goes, \u0026lsquo;a willow planted unintentionally grows into a shade tree.\u0026rsquo; The pneumonia detection project won a second prize in the South China region of C4-AI and successfully advanced to the national finals.\nAlso towards the end of this semester, I learned about the Mathematical Modeling Competition. I had previously known about the Mathematical Contest in Modeling, but I was simply too busy to review advanced calculus, so I let it go. I immediately decided to stay on campus over the summer for mathematical modeling training. However, after attending for just two days, I started studying on my own in the dormitory. One reason was that the instructor taught a bit too slowly, and we had to fight for seats in the large lecture hall. Another reason was that I wasn\u0026rsquo;t used to the option of attending lectures; I could completely self-study. So I took Jiang Qiyuan\u0026rsquo;s \u0026ldquo;Mathematical Models\u0026rdquo; and read through it in the dorm, though I only got a general overview. (To be honest, I feel that even if I had finished reading it, it wouldn\u0026rsquo;t have been very useful. The most important thing is doing the assignments and learning various models through that process. Since mathematical modeling is essentially an open-book exam, the modeling approach is paramount. Don\u0026rsquo;t aim too high while your skills are low; you can only discover certain details when you actually start modeling.)\nSophomore year ended in a rather perfunctory manner.\nConfusion The most important thing at the start of junior year was the Mathematical Modeling Competition. I teamed up with XM and WJ and worked for three days straight. Let me briefly summarize: On the first day, we received the problems, and by the first night, we had to decide which problem to tackle. Problem A was a physics problem; looking at it, we immediately ruled it out—I didn\u0026rsquo;t even read the full text. Problem B was an algorithm problem, likely NP-Hard. It felt like any random algorithm could solve it, though a more reliable approach would be Reinforcement Learning (RL). RL seemed capable of handling it, but this environment was too difficult, turning it into a test of programming skills. Although I knew RL could work, my understanding of it was limited to Professor Hongyi Li\u0026rsquo;s videos, and I had only skimmed the material without ever implementing it myself. My teammates couldn\u0026rsquo;t help much with programming; I was primarily responsible for coding, XM handled the main writing of the paper, and WJ knew some SPSS, MATLAB, and various mathematical models (he seemed to know more models than I did, possibly having attended more training courses). Using RL would definitely be innovative and likely to win an award. Problem C was a data processing problem, quite traditional. So I was torn between B and C, finally deciding to go back and try something first, then decide in the morning. That night, I coded a bit and reviewed RL, feeling it was feasible. The next morning, I went over and decided I would handle the RL part. Since the first question didn\u0026rsquo;t require RL at all—just the shortest path—I originally planned for one teammate to handle that while I tackled the RL for the subsequent questions, and another teammate would start writing the paper. After about 2 hours, the teammate working on the first question had no ideas. I thought, \u0026ldquo;Isn\u0026rsquo;t this just a casual write-up for C?\u0026rdquo; So I switched to writing the first question myself, which left one person idle. Finally, around 9 PM, we decided to choose Problem C. The challenge with C was finding an innovative point. However, since C was all about data processing, I used Jupyter Notebook and messed around with pandas. Fortunately, I had done correlation analysis in a previous assignment (an analysis of Earth\u0026rsquo;s surface temperature, likely a past exam question), so data processing posed no real difficulties. I even used machine learning for classification, primarily using decision trees (since they performed well). On the last night, there seemed to be a major issue with the approach for one question, and I spent a long time thinking of a solution. I slept only about 3 hours that night. I rushed through most of the work, went back to rest for 2 hours at noon, then returned in the afternoon to supplement and polish the paper, organizing the final data and answers. However, right before submission, I realized that the data I had previously given to one teammate had not been updated. I didn\u0026rsquo;t know what happened, so I hurriedly asked him to modify it. It was obvious—no, I could clearly see—that his hands were shaking violently. I then fixed this part myself, completing the verification almost at the deadline, and waited for submission. The competition was over. One notable thing was that we worked at Siyuan, where most people were studying. To avoid disturbing the neighbors, we did our mathematical modeling in a place without air conditioning and no one else studying. I was soaked with sweat, while my teammates had already become \u0026ldquo;Buddhas\u0026rdquo; (calm and detached) hhh. I guess I just got used to it. Only on the last night, after everyone else had left, did we move to a place with air conditioning for the night, then moved back early the next morning.\nThe final result was gratifying: a National Second Prize and a Provincial First Prize.\nThen came the usual recruitment drive. I can\u0026rsquo;t remember other details; it was just a bunch of miscellaneous small matters. Meanwhile, my major courses increased, bringing significant academic pressure. My focus became scattered, and I didn\u0026rsquo;t review as seriously for the finals as before. I adopted a \u0026lsquo;one subject per day\u0026rsquo; approach, taking exams in a laid-back, \u0026lsquo;Buddha-like\u0026rsquo; manner.\nBy the way, of course there\u0026rsquo;s also the national competition I mentioned earlier—a fully funded trip to Hangzhou. I only won a third prize (x), but it did reveal many issues with our project. I felt a bit unwilling to lose because the judges immediately asked, \u0026ldquo;Do you know what smart healthcare is?\u0026rdquo; That question really caught me off guard; it was completely unexpected. I felt our project name was too grand, our focus too narrow, and our work insufficient. Additionally, another impact of this competition was that the day after it ended, I had the CET-6 exam. I essentially took it completely unprepared, so the result was predictable, but the impact was significant since this was the final exam before the recommendation for graduate studies.\nThere was also the Tiantie competition. With RK\u0026rsquo;s help, we assembled a \u0026lsquo;Galaxy Battleship\u0026rsquo; team, easily securing a provincial first prize. I personally won a national second prize. I only found out about this competition this year; otherwise, I would have participated in my freshman year (though I\u0026rsquo;m not sure if anyone would have been willing to take me along).\nAnother small competition was the Maker Marathon. I handled the full-stack content myself, including data labeling, working on the \u0026lsquo;Sky Brush\u0026rsquo; project ([code], [demo]). We won a school first prize, and it was also my first exposure to the field of object detection.\nBy the second semester of my junior year, the two previous projects that hadn\u0026rsquo;t yielded results continued. I took on the role of lead for the FL pneumonia detection project. After President Luo arrived, the university established the School of Biomedical Engineering, which aligned well with this project. I found a professor from that school to serve as an advisor. Of course, I was the project lead, and the project belonged to the School of Biomedical Engineering, yet I am not a student of that school. So, I can only say our college lost out (x).\nWith ZW\u0026rsquo;s help, I polished this project. However, we stumbled in our first attempt: we won a school first prize in the Challenge Cup, but only a provincial third prize. Later, as the project lead, I felt somewhat irresponsible, as most of the tasks were actually handled by ZW, ZC, and SY.\nAround this time, my grades had been declining year by year, dropping to rank 11. I felt my chances for recommendation were hopeless or at least very slim (since our school\u0026rsquo;s policy allows a maximum of 0.3 points for awards and papers; I could easily max that out, but others might not. Based on the 5% recommendation rate in previous years, I probably needed to be ranked 9th. I felt there was still a slight chance, but not a strong one—it\u0026rsquo;s that edge-case dilemma, you know?). So, I decided to apply to universities in Hong Kong. My parents were very supportive, so I started studying for the IELTS. I read many articles on Zhihu, but most were marketing fluff; I\u0026rsquo;d advise against reading them as their reference value is low since everyone\u0026rsquo;s foundation and aptitude differ. At the time, I just studied on my own, casually memorizing vocabulary and doing practice questions. However, while reviewing for finals, I forgot to register for the first IELTS exam in the summer. G, so I couldn\u0026rsquo;t take the exam until the end of July.\nDuring this period, I contacted Professor W at a Hong Kong university for a brief interview and presentation. Since I had no experience, I didn\u0026rsquo;t know what to expect in an application interview. The questions covered linear algebra, advanced calculus, and probability, and I performed terribly—I was even embarrassed by my own answers. In the end, I didn\u0026rsquo;t receive an offer. However, I asked the professor if there were any opportunities to work on projects or other tasks, mainly hoping to produce results for a paper. The professor assigned me a project, which became my main activity over the summer besides preparing for the IELTS.\nJoy How did the summer IELTS go? 14 and then calmly signed up for a class. It wasn\u0026rsquo;t ideal, but things proceeded steadily.\nWhen senior year began, the recommendation process started. With a glimmer of hope, I still submitted my materials. Many Hong Kong universities had already opened their early application rounds. Since Hong Kong universities require language scores before applying, and it\u0026rsquo;s first-come, first-served, I was quite panicked without any language scores. Additionally, at the end of the previous semester, I tried out for summer camps at a few schools, but unfortunately, none accepted me. I think there were a few reasons: 1. I hadn\u0026rsquo;t passed the CET-6, which is very important for recommendations; regardless of the score, one should at least pass. 2. The bar was too high, and my undergraduate university was a limiting factor. Fortunately, I passed the CET-6 in the previous semester with a score of 478, and that score was used for the recommendation.\nFinally, when I saw that most people could max out the 0.3 points, I felt panicked. However, when the comprehensive scores were released, I was ranked 8th, which meant I had a very good chance. I quickly contacted professors and indeed disturbed many of them, applying to pre-recommendation programs at various schools. With my CET-6 score, the process was obviously much smoother. When the final recommendation quotas were announced, the number of spots increased significantly to 11, which was a sure thing.\nAround the evening of September 13th, I received a call from a lab at the Institute of Computing Technology (ICT), though it wasn\u0026rsquo;t one I had applied to. They asked if I was willing to go, mentioning there might be a written test and an interview, with specific details to be announced the next day. Of course, I happily agreed. The next day, I received an interview notice from the lab I had actually applied to, which was different from the previous one. I started preparing for the nucleic acid test and the interview. Because of that earlier call, I thought there would always be a written test, so I was quite panicked and arrived a day early. I had intended to book a hotel near the ICT, but ended up booking one directly opposite the institute\u0026hellip; It was too close\u0026hellip; But you can imagine that because of this, the hotel room rate nearly doubled; I recall it was 300+ RMB per night.\nHowever, I never received a written test notice later on, leaving me completely confused. I attended the lab interview and that was it. The interview process went relatively smoothly. At first, they asked for a self-introduction. I had read many interview guides on Zhihu, which suggested starting with an English introduction, followed by simple English questions, then technical questions, and finally project-related questions. However, when I got there, the professor just said, \u0026ldquo;Introduce yourself.\u0026rdquo; I hadn\u0026rsquo;t even had the courage to pull out my prepared English introduction, so I ended up translating my English draft back into Chinese\u0026hellip; Then, I hadn\u0026rsquo;t had time to review the technical questions, so they just asked some questions based on my resume.\nThree interludes: Before receiving the interview notice from the Institute of Computing Technology, I actually received a phone interview, which gave me a rough idea of potential issues in my resume. Another was an interview with a professor from Xiamen University (XMU) I contacted before the Institute interview; it happened on the morning of the Institute interview day. The professor was quite welcoming, and honestly, I quite liked that professor hhh. These interview experiences helped me stay calmer and more confident in future interviews. The third was spending an afternoon sightseeing in Beijing hhh. Let me complain: Beijing 7-Elevens actually have no seats for eating\u0026hellip; Thinking back to convenience stores in Haikou, they all have seats\u0026hellip; So I just sat by the window and ate something.\nAfter the interview, I went to celebrate. Since the Mid-Autumn Festival was mostly a holiday, and few universities organize pre-recommendation interviews during this period, I decided to play for two days before returning to continue applications. That night, while walking in Nanluoguxiang, coincidentally, I had originally wanted to visit \u0026lsquo;Little Bear Diary\u0026rsquo; but couldn\u0026rsquo;t find it on the map for ages. Yet, as I walked, I suddenly spotted it. Just as I stepped inside, I received an offer! So I immediately bought a Little Bear plushie, named it \u0026lsquo;Offer\u0026rsquo;, abbreviated as \u0026lsquo;Little O!\u0026rsquo; With the pressure gone, I naturally felt much more relaxed. On the way back to the hotel that night, I found a print shop and signed the agreement.\nAfter that, I declined all subsequent interviews and pre-recommendation processes, and informed all the professors I had previously contacted about the situation.\nThen I waited for the recommendation system to open. There were two waits: one was for Hainan University to process my application, which they submitted on the 25th, and I could only check my recommendation qualification on the evening of the 26th. The other was waiting for the admission notice from the Institute of Computing Technology after registration.\nDuring the wait for the Institute\u0026rsquo;s notice, I was fortunate to join the \u0026lsquo;Green Group\u0026rsquo; and met many notable figures, such as Brother Ocean! Brother Ocean is truly inspirational! Of course, the Institute also produced some iconic moments. The long wait was always agonizing. For two days with no notification, I had no mood to do anything else, mostly fearing unexpected outcomes. Meanwhile, the various events after September 28th were also emotional. Some people held onto too many offers and only released them during the system registration, causing schools to be massively ghosted, forced to fill spots with late admissions, or even completely ghosted, requiring them to prepare for a second pre-recommendation round. In my view, schools that over-issue offers are even worse than students who play the field.\nWhen a school gets ghosted, at least they have backups. But if a student accepts an offer without keeping any backups and then gets ghosted by the school, they truly have no way out and can only hope to pick up leftovers. I believe the most important thing during this period is honesty. I can\u0026rsquo;t help but complain: for those \u0026lsquo;sea kings\u0026rsquo; (players), why not release offers from schools you know you won\u0026rsquo;t attend? Wouldn\u0026rsquo;t it be better for schools to fill spots early? As for schools that over-issue offers, I have nothing to say\u0026hellip; They deserve their ruined reputations. This vicious cycle is: students keep multiple offers (who knows which school will ghost them, after all, having a place to study is most important), and school admissions staff can only create long waitlists. Schools that over-issue offers will inevitably get ghosted by many. Ultimately, true recommendations still depend on September 28th, xs. Of course, we can\u0026rsquo;t guarantee that no one will ghost anyone, or that no school will ghost students,after all, school agreements are just pieces of paper.\nThank you, Institute of Computing Technology, for taking me in. I recommend everyone apply to the Institute\u0026rsquo;s 2023 cohort, even if it\u0026rsquo;s a bit too safe, but it counts as my first graduate school lesson hhhh. I received the admission notice at 11 PM on September 30th, by which time I had already started eating at Haidilao 2333.\nDuring the waiting period, a professor from Hainan University also called me, saying that if I wasn\u0026rsquo;t admitted elsewhere, they would be willing to let me pursue my master\u0026rsquo;s at my home university and would provide the best faculty in the college. I was truly moved 555.\nThe Aftermath During my four years of university, I had two romantic relationships, and there will only be two.\nThis article might not have many pictures. If I remember later, I\u0026rsquo;ll come back to add some, mainly because it\u0026rsquo;s troublesome.\nUnfinished Business Whether known or unknown, it doesn\u0026rsquo;t matter that much. After all, everything has already happened, so let the story continue.\nI actually wanted to write this memoir back in high school, but first, due to psychological shame (probably because the outcome wasn\u0026rsquo;t great), and second, because I didn\u0026rsquo;t want to recall it (probably a form of avoidance).\nThis piece does feel a bit like a diary entry, but not entirely. I just wanted to share my experience with you.\nI may have never shared this with anyone except my closest loved ones. Now, there\u0026rsquo;s you.\nThat\u0026rsquo;s enough, I\u0026rsquo;m getting emo. This article is nearly 10,000 words.\nWritten on October 2, 2021, at 4:26 AM.\nAcknowledgements Thank you to my parents for raising me, for not bringing family pressure upon me, allowing me more opportunities to find the direction I pursue. Most importantly, thank you for respecting my opinions and choices.\nThank you to my darling who has always been by my side, encouraging me.\nThank you to all the teachers who nurtured me along the way, whether they were course instructors or professors I contacted during my recommendation or application process. Getting to know you all has been a wealth for me, and I am especially grateful that the teachers I encountered were so gentle and kind.\nThank you to the motherland and my university for providing a good living environment.\nThank you to Senior BY for my research enlightenment and guidance.\nThank you to the friends who encouraged each other during my low points, whether roommates or project partners. Maybe I wasn\u0026rsquo;t always the best person.\nAbout Me Here\u0026rsquo;s a contact method: you can find my contact information on the About page, or leave a message below this blog post.\n1\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n2\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n3\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nArduino\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHDCTF\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n1\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n2\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nHTML\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nJavaScript\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n1\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n2\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRaspberry Pi for MNIST\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nRaspberry Pi for Fire Recognition\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nTook one exam\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://blog.bj-yan.top/en/p/journey-man-man-qiu-xue-lu/","summary":"\u003ch2 id=\"preface\"\u003ePreface\u003c/h2\u003e\n\u003cp\u003eSeptember 30, 2021. Finally, the weight on my heart has been lifted.\u003c/p\u003e\n\u003cp\u003eAlthough I had long known the outcome was a done deal, only when I clicked \u0026lsquo;Accept\u0026rsquo; in the system did I feel that everything had finally settled.\u003c/p\u003e\n\u003cp\u003eThis is already my fourth year of university life. From that initially withdrawn, introverted, and taciturn high school student, I have gradually grown into who I am today.\u003c/p\u003e\n\u003cp\u003eOnly now, feeling that my efforts have finally yielded corresponding rewards, do I feel that everything was worth it.\u003c/p\u003e","title":"A Long Journey of Learning"},{"content":"Just here to brainstorm. Problem A looks impossible at first glance—it\u0026rsquo;s a physics problem; Problem B seems quite simple, though I haven\u0026rsquo;t looked into it closely; Problem C doesn\u0026rsquo;t seem too difficult either.\nSo I\u0026rsquo;ll mainly write up the approach for Problem C. Since I\u0026rsquo;m not actually solving it myself, some details might not be fully considered; this is just one possible feasible solution.\nProblem C: Raw Material Procurement and Transportation for a Manufacturing Enterprise 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 C 题 生产企业原材料的订购与运输 某建筑和装饰板材的生产企业所用原材料主要是木质纤维和其他植物素纤维材料, 总体可分为 A，B，C 三种类型。该企业每年按 48 周安排生产，需要提前制定 24 周的原 材料订购和转运计划，即根据产能要求确定需要订购的原材料供应商（称为“供应商”） 和相应每周的原材料订购数量（称为“订货量”），确定第三方物流公司（称为“转运 商”）并委托其将供应商每周的原材料供货数量（称为“供货量”）转运到企业仓库。 该企业每周的产能为 2.82 万立方米，每立方米产品需消耗 A 类原材料 0.6 立方米， 或 B 类原材料 0.66 立方米，或 C 类原材料 0.72 立方米。由于原材料的特殊性，供应商 不能保证严格按订货量供货，实际供货量可能多于或少于订货量。为了保证正常生产的 需要，该企业要尽可能保持不少于满足两周生产需求的原材料库存量，为此该企业对供 应商实际提供的原材料总是全部收购。 在实际转运过程中，原材料会有一定的损耗（损耗量占供货量的百分比称为“损耗 率”），转运商实际运送到企业仓库的原材料数量称为“接收量”。每家转运商的运输 能力为 6000 立方米/周。通常情况下，一家供应商每周供应的原材料尽量由一家转运商 运输。 原材料的采购成本直接影响到企业的生产效益，实际中 A 类和 B 类原材料的采购单 价分别比 C 类原材料高 20%和 10%。三类原材料运输和储存的单位费用相同。 附件 1 给出了该企业近 5 年 402 家原材料供应商的订货量和供货量数据。附件 2 给 出了 8 家转运商的运输损耗率数据。请你们团队结合实际情况，对相关数据进行深入分 析，研究下列问题： 1．根据附件 1，对 402 家供应商的供货特征进行量化分析，建立反映保障企业生产 重要性的数学模型，在此基础上确定 50 家最重要的供应商，并在论文中列表给出结果。 2．参考问题 1，该企业应至少选择多少家供应商供应原材料才可能满足生产的需求？ 针对这些供应商，为该企业制定未来 24 周每周最经济的原材料订购方案，并据此制定 损耗最少的转运方案。试对订购方案和转运方案的实施效果进行分析。 3．该企业为了压缩生产成本，现计划尽量多地采购 A 类和尽量少地采购 C 类原材 料，以减少转运及仓储的成本，同时希望转运商的转运损耗率尽量少。请制定新的订购 方案及转运方案，并分析方案的实施效果。 4．该企业通过技术改造已具备了提高产能的潜力。根据现有原材料的供应商和转运 商的实际情况，确定该企业每周的产能可以提高多少，并给出未来 24 周的订购和转运 方案。 注：请将问题 2、问题 3 和问题 4 订购方案的数值结果填入附件 A，转运方案的数 值结果填入附件 B，并作为支撑材料（勿改变文件名）随论文一起提交。 附件 1 的数据说明 （1）企业的订货量：第一列为供应商的名称；第二列为供应商供应原材料的类别； 第三列及以后共 240 列为企业向各供应商每周的订货量（单位：立方米）；数值“0”表 示相应的周（所在列）没有向供应商（所在行）订货。 （2）供应商的供货量：第一列为供应商的名称；第二列为供应商供应原材料的类别； 第三列及以后共 240 列为各供应商每周的供货量（单位：立方米）；数值“0”表示相应 的周（所在列）供应商（所在行）没有供货。 附件 2 的数据说明 第一列为转运商的名称；第二列及以后共 240 列为每周各转运商的运输损耗率（%）， 即 损耗率 = (供货量 - 接收量) / (供货量) X 100%；数值“0”表示没有运送。 Problem 1 Based on Attachment 1, perform a quantitative analysis of the supply characteristics of the 402 suppliers. Establish a mathematical model reflecting the importance of each supplier in ensuring the enterprise\u0026rsquo;s production. Based on this, identify the 50 most critical suppliers and present the results in a table within the paper. The first problem is clearly an evaluation model, and the result will likely be used repeatedly later. First, determine evaluation indicators, such as delivery completion rate, total supply volume, and supply trends. You can use gray prediction or fuzzy evaluation models to set weights for each indicator, or search for relevant articles on this type of evaluation. Finally, provide a ranked table.\nProblem 2 Referencing Problem 1, how many suppliers should the enterprise select at minimum to potentially meet production demands? For these suppliers, formulate the most economical weekly raw material procurement plan for the next 24 weeks, and based on this, develop a transportation plan with the least loss. Analyze the implementation effectiveness of both the procurement and transportation plans.\nThere is a very important condition in the problem statement为了保证正常生产的需要，该企业要尽可能保持不少于满足两周生产需求的原材料库存量，为此该企业对供应商实际提供的原材料总是全部收购。. I didn\u0026rsquo;t notice this condition at first, but when I reconsidered it, there was a significant change.\nThis question mainly involves two issues: transportation and supply.\nFirst, let\u0026rsquo;s answer a question from the problem statement; it\u0026rsquo;s quite simple. However, before solving all subsequent problems, we need to predict the supply volume of the 50 suppliers identified in Problem 1 for the next 24 weeks. There are many methods available; the simplest is linear regression. Once we have this, sort by supply volume, sum from largest to smallest until the supply for 2 weeks is covered. This can be solved via brute force or binary search; given the small data size, a simple iteration will suffice.\nThe most economical procurement plan and transportation plan are relatively independent problems, but they must be solved in a specific order. First, solve the transportation plan. For each transporter, calculating an average loss rate should be sufficient, or use other metrics to compute and rank them, then select from the top down. For the procurement plan, prioritize suppliers with high completion rates and the lowest cost after deducting loss rates and utilization rates. I\u0026rsquo;d guess the priority raw material type would be A.\nProblem 3 To reduce production costs, the enterprise plans to procure as much of raw material type A as possible and as little of type C as possible, thereby reducing transportation and warehousing costs, while also hoping to minimize the loss rates of transporters. Please formulate new procurement and transportation plans and analyze the implementation effectiveness of these plans.\nThis is an optimization problem and feels quite open-ended. First, establish a formula for total or comprehensive cost. Then, apply optimization methods such as Newton\u0026rsquo;s iteration, linear programming, simulated annealing, hill climbing, or ant colony optimization to find a solution. There is no standard answer, of course, but the better the optimization, the better.\nProblem 4 Through technical transformation, the enterprise has now acquired the potential to increase production capacity. Based on the current actual conditions of raw material suppliers and transporters, determine how much the weekly production capacity can be increased, and provide the procurement and transportation plans for the next 24 weeks.\nThe increase in production capacity is constrained by two factors: supply limitations and transportation limitations. For supply constraints, the strategy for the initial weeks should be to transport as much as possible to set the stage for later, but this is limited by transportation capacity. Therefore, the weekly increase should be maximized under other conditions, capped at the upper limit of the constraints.\nThinking back to the last time I participated in the Mathematical Modeling Competition, it feels like just yesterday. Time flies, and things have changed. The senior student who competed with me last year didn\u0026rsquo;t pursue graduate studies but went straight to work, while another classmate of mine is expected to have successfully secured a recommendation for graduate school and already holds three job offers orz I\u0026rsquo;ve also transitioned from a contestant to a contestant who can freely ramble (x)\n","permalink":"https://blog.bj-yan.top/en/p/blog-mum-2021/","summary":"\u003cp\u003eJust here to brainstorm. Problem A looks impossible at first glance—it\u0026rsquo;s a physics problem; Problem B seems quite simple, though I haven\u0026rsquo;t looked into it closely; Problem C doesn\u0026rsquo;t seem too difficult either.\u003c/p\u003e\n\u003cp\u003eSo I\u0026rsquo;ll mainly write up the approach for Problem C. Since I\u0026rsquo;m not actually solving it myself, some details might not be fully considered; this is just one possible feasible solution.\u003c/p\u003e\n\u003ch2 id=\"problem-c-raw-material-procurement-and-transportation-for-a-manufacturing-enterprise\"\u003eProblem C: Raw Material Procurement and Transportation for a Manufacturing Enterprise\u003c/h2\u003e\n\u003cdiv class=\"highlight\"\u003e\u003cdiv class=\"chroma\"\u003e\n\u003ctable class=\"lntable\"\u003e\u003ctr\u003e\u003ctd class=\"lntd\"\u003e\n\u003cpre tabindex=\"0\" class=\"chroma\"\u003e\u003ccode\u003e\u003cspan class=\"lnt\"\u003e 1\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e 2\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e 3\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e 4\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e 5\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e 6\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e 7\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e 8\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e 9\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e10\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e11\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e12\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e13\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e14\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e15\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e16\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e17\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e18\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e19\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e20\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e21\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e22\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e23\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e24\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e25\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e26\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e27\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e28\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e29\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e30\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e31\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e32\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e33\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e34\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e35\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e36\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e37\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e38\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e39\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e40\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e41\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e42\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e43\n\u003c/span\u003e\u003cspan class=\"lnt\"\u003e44\n\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\u003c/td\u003e\n\u003ctd class=\"lntd\"\u003e\n\u003cpre tabindex=\"0\" class=\"chroma\"\u003e\u003ccode class=\"language-text\" data-lang=\"text\"\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003eC 题 生产企业原材料的订购与运输\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e某建筑和装饰板材的生产企业所用原材料主要是木质纤维和其他植物素纤维材料,\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e总体可分为 A，B，C 三种类型。该企业每年按 48 周安排生产，需要提前制定 24 周的原\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e材料订购和转运计划，即根据产能要求确定需要订购的原材料供应商（称为“供应商”）\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e和相应每周的原材料订购数量（称为“订货量”），确定第三方物流公司（称为“转运\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e商”）并委托其将供应商每周的原材料供货数量（称为“供货量”）转运到企业仓库。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e该企业每周的产能为 2.82 万立方米，每立方米产品需消耗 A 类原材料 0.6 立方米，\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e或 B 类原材料 0.66 立方米，或 C 类原材料 0.72 立方米。由于原材料的特殊性，供应商\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e不能保证严格按订货量供货，实际供货量可能多于或少于订货量。为了保证正常生产的\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e需要，该企业要尽可能保持不少于满足两周生产需求的原材料库存量，为此该企业对供\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e应商实际提供的原材料总是全部收购。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e在实际转运过程中，原材料会有一定的损耗（损耗量占供货量的百分比称为“损耗\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e率”），转运商实际运送到企业仓库的原材料数量称为“接收量”。每家转运商的运输\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e能力为 6000 立方米/周。通常情况下，一家供应商每周供应的原材料尽量由一家转运商\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e运输。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e原材料的采购成本直接影响到企业的生产效益，实际中 A 类和 B 类原材料的采购单\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e价分别比 C 类原材料高 20%和 10%。三类原材料运输和储存的单位费用相同。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e附件 1 给出了该企业近 5 年 402 家原材料供应商的订货量和供货量数据。附件 2 给\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e出了 8 家转运商的运输损耗率数据。请你们团队结合实际情况，对相关数据进行深入分\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e析，研究下列问题：\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e1．根据附件 1，对 402 家供应商的供货特征进行量化分析，建立反映保障企业生产\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e重要性的数学模型，在此基础上确定 50 家最重要的供应商，并在论文中列表给出结果。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e2．参考问题 1，该企业应至少选择多少家供应商供应原材料才可能满足生产的需求？\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e针对这些供应商，为该企业制定未来 24 周每周最经济的原材料订购方案，并据此制定\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e损耗最少的转运方案。试对订购方案和转运方案的实施效果进行分析。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e3．该企业为了压缩生产成本，现计划尽量多地采购 A 类和尽量少地采购 C 类原材\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e料，以减少转运及仓储的成本，同时希望转运商的转运损耗率尽量少。请制定新的订购\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e方案及转运方案，并分析方案的实施效果。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e4．该企业通过技术改造已具备了提高产能的潜力。根据现有原材料的供应商和转运\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e商的实际情况，确定该企业每周的产能可以提高多少，并给出未来 24 周的订购和转运\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e方案。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e注：请将问题 2、问题 3 和问题 4 订购方案的数值结果填入附件 A，转运方案的数\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e值结果填入附件 B，并作为支撑材料（勿改变文件名）随论文一起提交。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e附件 1 的数据说明\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e（1）企业的订货量：第一列为供应商的名称；第二列为供应商供应原材料的类别；\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e第三列及以后共 240 列为企业向各供应商每周的订货量（单位：立方米）；数值“0”表\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e示相应的周（所在列）没有向供应商（所在行）订货。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e（2）供应商的供货量：第一列为供应商的名称；第二列为供应商供应原材料的类别；\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e第三列及以后共 240 列为各供应商每周的供货量（单位：立方米）；数值“0”表示相应\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e的周（所在列）供应商（所在行）没有供货。\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e附件 2 的数据说明\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e第一列为转运商的名称；第二列及以后共 240 列为每周各转运商的运输损耗率（%），\n\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"line\"\u003e\u003cspan class=\"cl\"\u003e即 损耗率 = (供货量 - 接收量) / (供货量) X 100%；数值“0”表示没有运送。\n\u003c/span\u003e\u003c/span\u003e\u003c/code\u003e\u003c/pre\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\u003ch3 id=\"problem-1\"\u003eProblem 1\u003c/h3\u003e\n\u003cblockquote\u003e\n\u003col\u003e\n\u003cli\u003eBased on Attachment 1, perform a quantitative analysis of the supply characteristics of the 402 suppliers. Establish a mathematical model reflecting the importance of each supplier in ensuring the enterprise\u0026rsquo;s production. Based on this, identify the 50 most critical suppliers and present the results in a table within the paper.\u003c/li\u003e\n\u003c/ol\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eThe first problem is clearly an evaluation model, and the result will likely be used repeatedly later. First, determine evaluation indicators, such as delivery completion rate, total supply volume, and supply trends. You can use gray prediction or fuzzy evaluation models to set weights for each indicator, or search for relevant articles on this type of evaluation. Finally, provide a ranked table.\u003c/p\u003e","title":"2021 Mathematical Modeling"},{"content":"Before I originally wanted to sign up for the first July session early to get a feel for the test, but it just so happened to coincide with exam week, so I missed the registration deadline. No choice but to sign up for the next session, which is this one.\nHainan University does not have a computer-based testing center, so I had to take the paper-based exam. The venue was in the teaching building inside Building 5.\nIt\u0026rsquo;s 13:03 now. I just finished the exam, played a round of Teamfight Tactics, and had a meal. Let\u0026rsquo;s write this stream-of-consciousness account while it\u0026rsquo;s still fresh.\nSpeaking My Speaking test was scheduled for the 23rd at 13:50. I rushed through lunch, printed my admission ticket, and headed over. I arrived around 13:10. To be honest, I was extremely panicked and had no idea what to expect.\nAfter arriving at the test center, they did the usual checks for access codes, health codes, and vaccination status, then immediately ushered me to the waiting room to drop off my belongings.\nI told the guy I wanted to check my phone a bit more, but he said that wasn\u0026rsquo;t allowed in the waiting room. I asked if I could go outside, and he said yes. So I stood outside and checked my phone. However, the proctor inside told me to go to the waiting room at the very end, which is used for taking photos and identity verification. So I went there and waited until just after 13:20 before they called me.\nIdentity verification involved taking off my mask and glasses, recording my fingerprints four times (right index finger), and taking a photo. Then I waited. There were two others with me, labeled S1, S2, and S3; I was S1.\nWhen it was almost time, another proctor escorted people to the fourth floor, one person per classroom. Before leaving, they did another identity verification, checking fingerprints again and verifying my face by removing my mask and glasses.\nWhen I got to the fourth floor, the examiner was already standing at the door. I went over, closed the door behind me, and was led to my seat.\nThe seats were separated by a transparent partition in the middle, probably for pandemic prevention. Then the exam process began.\nPart1 They simply asked for my full name and whether I had brought any devices.\nPart2 For Part 2, they gave me a pencil, which was already placed on the right side of the desk, along with a piece of paper for taking notes. The topic was whether I often share with people, when I first started sharing, and with whom. I vaguely remember my answer was a complete mess. I said I share machine learning knowledge with classmates from my association during online meetings. One of my points was that for any knowledge, when you learn to share it with others so they can understand it, that\u0026rsquo;s when you truly master it. Although the general meaning was there, I\u0026rsquo;m sure my expression and language organization were terrible. (I really don\u0026rsquo;t want to recall this part.) Later, because I hadn\u0026rsquo;t spoken for the full 2 minutes, the examiner asked, \u0026lsquo;Could you tell something more?\u0026rsquo; I stammered for a while but couldn\u0026rsquo;t figure out what to say.\nPart3 I remember one question was about why I don\u0026rsquo;t share with people. I said \u0026lsquo;secrets?\u0026rsquo; Then, in a clumsy attempt to add more, I said no one wants to make friends with people who don\u0026rsquo;t share, and those people are selfish. The examiner said people who don\u0026rsquo;t share can still make friends. I was stunned and then forgot the rest.\nFinally, the examiner said the test was over and escorted me out. It was past 14:00 when I came out, and I was handed an IELTS pencil. The guy in charge of item storage hadn\u0026rsquo;t given me my wristband yet; he seemed a bit weird.\nThe paper-based exam was the next morning. I originally planned to wake up at 6:30 to review essay materials, but\u0026hellip; 6:30, 7:00, 7:20—I set three alarms, and not a single one rang??? I woke up at 7:22. I was dizzy. Then I went straight to the exam without eating breakfast. After stepping outside, I realized I hadn\u0026rsquo;t brought my mask, so I had to go back.\nWhen I arrived, I was incredibly thirsty. I hadn\u0026rsquo;t drunk any water before leaving, so I planned to buy a bottle from the vending machine downstairs in Building 5, but the vending machine was turned off\u0026hellip; Once I entered the exam area, it was the usual routine. I waited in the waiting room. Fortunately, the waiting room had small bottles of Nongfu Spring mineral water, but I had to tear off the labels. It turned out that for these two days, I didn\u0026rsquo;t actually need to print my admission ticket; just bringing my ID card and knowing the Speaking test time was enough.\nBefore entering the exam room, there was the usual security check. Inside, I had to wash my hands with waterless sanitizer, sign a document (this part was in English), and after entering the exam room, everything seemed to be in English.\nListening The first two Listening sections were okay; I knew a few words but forgot how to spell them, so I lost points for free. One was \u0026lsquo;balcony\u0026rsquo; Balcony, and the other was \u0026lsquo;refrigerator\u0026rsquo; Refrigerator.\nLater, there was a section where I had to match content. I didn\u0026rsquo;t understand the whole Listening section at all; there were too many synonym replacements. I need to review this part again. I felt the last two Listening sections were quite difficult (for me).\nListening took 30 minutes, with the last 10 minutes allocated for transferring answers.\nReading I hadn\u0026rsquo;t really timed my Reading sections before. Normally, each passage should take about 20 minutes, but the first one took me nearly 25 minutes. Mainly, I didn\u0026rsquo;t realize there were so many questions at the beginning, and finding answers one by one was slow. This directly caused me to spend only about 10 minutes on the last passage: 5 minutes for answering questions. Fortunately, it was a sentence completion section. I found the answers for the following multiple-choice questions in the last two paragraphs, but for the matching questions in the middle, I simply ran out of time. When they were about to collect the papers, I had to guess the answers BCDEF. I just had to leave the rest to fate, lol.\nWriting I felt the Task 1 (small essay) was relatively simple. It had two pie charts showing water usage in various aspects in Sydney, Australia. I had seen similar ones before, so I was okay with it. Personally, I felt good about it. The requirement was 150 words, and I just wrote casually and reached that count.\nThe essay task was a bit strange: it presented some people think that in modern world people are becoming more dependent on each other, then another viewpoint was independent, and the requirement was discuss both view followed by asking for my own opinion.\nI remember I roughly outlined it, but I don\u0026rsquo;t think I reached 250 words. Later, I felt time was insufficient; I spent 20 minutes on ideas and 20 minutes on writing content, which was indeed an unreasonable allocation of time.\nIntroduction: background + two viewpoints Viewpoint A: independent. In the past, people lived in villages, rising with the sun and resting at sunset, hunting and raising animals together (I couldn\u0026rsquo;t write about farming, md). Nowadays, people live in apartments or houses with different lifestyles. Students leave school at 7:00 AM, office workers head to work at 8:00 AM, and they don\u0026rsquo;t even have time to say a few words. After school, students still have homework to do, and office workers have leftover work to finish. Viewpoint B: dependent. Many people are willing to help others when in the truble, or donate, so dependent. My argument: I thought we should mental on dependent and action on dependent, because we live in the same world and cannot be completely independent, and then\u0026hellip;\nI don\u0026rsquo;t know if I went off-topic\u0026hellip;\nEnd Finally, make a proper summary.\nI should probably make a review plan soon, and also find a one-on-one class\u0026hellip;\nWhy call it a one-day trip when it was two days? Because the speaking test was at 13:50 in the afternoon, so indeed everything was completed within a single day - -\n","permalink":"https://blog.bj-yan.top/en/p/misc-20210724/","summary":"\u003ch2 id=\"before\"\u003eBefore\u003c/h2\u003e\n\u003cp\u003eI originally wanted to sign up for the first July session early to get a feel for the test, but it just so happened to coincide with exam week, so I missed the registration deadline. No choice but to sign up for the next session, which is this one.\u003c/p\u003e\n\u003cp\u003eHainan University does not have a computer-based testing center, so I had to take the paper-based exam. The venue was in the teaching building inside Building 5.\u003c/p\u003e","title":"A Day Trip to My First IELTS Exam"},{"content":"Hello, welcome to this page, thank you for visiting.\nI am Xiao Bing.\nAn ordinary college student.\nEnergy Thief \u0026amp; Level 4 Chicken Pummeler I will update some essays or knowledge sharing here, hoping it helps you.\nI quite like a quote from Changzi:\nJust remember me casually, then forget me.\nMy academic homepage is here: link\nAlso, my resume [ZH] [EN]\nOf course, you can also leave a message here; I will check it from time to time. You can also directly add my WeChat to discuss with me.\nWeChat: yyyanbj。\nCriticism is welcome. Ways to criticize:\nEmail: bj.yan.pa@qq.com bj.yan@ieee.org\n","permalink":"https://blog.bj-yan.top/en/about/","summary":"\u003cp\u003eHello, welcome to this page, thank you for visiting.\u003c/p\u003e\n\u003cp\u003eI am Xiao Bing.\u003c/p\u003e\n\u003cp\u003eAn ordinary college student.\u003c/p\u003e\n\u003cp\u003e\u003cdel\u003e Energy Thief \u0026amp; Level 4 Chicken Pummeler \u003c/del\u003e\u003c/p\u003e\n\u003cp\u003eI will update some essays or knowledge sharing here, hoping it helps you.\u003c/p\u003e\n\u003cp\u003eI quite like a quote from Changzi:\u003c/p\u003e\n\u003cblockquote\u003e\n\u003cp\u003eJust remember me casually, then forget me.\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003cp\u003eMy academic homepage is here: \u003ca href=\"https://www.bj-yan.top/\"\u003elink\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eAlso, my resume \u003ca href=\"https://yyyanbj.github.io/pdf/cv_cn.pdf\"\u003e[ZH]\u003c/a\u003e \u003ca href=\"https://yyyanbj.github.io/pdf/cv_en.pdf\"\u003e[EN]\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003eOf course, you can also leave a message here; I will check it from time to time. You can also directly add my WeChat to discuss with me.\u003c/p\u003e","title":"About me"},{"content":" Occasionally I check in on my friends\u0026rsquo; friend links. If some old friends\u0026rsquo; domains or websites have become invalid, they may be temporarily removed. Please leave a message with your new site information~ (The following list is not ranked)\nYuZhangWang 的领域 PLUSULTRA!\nRussell Keep eating codes!\nNelson Boss Swim until the sea turns blue\nJunyi\u0026#39;s Lab A place for self-amusement\n月梦の技术博客 Learn broadly, hold fast to your aspirations; ask earnestly, reflect on what is near.\n孑渡 VR Big Brother.\nVnYzm On the edge of a remote small fishing village\nChasing1020 Why there is a universe?\nR0gerThat @Vidar-Team @CTFer @Game Lover\nDimsmary Do What U Wanna Do\nDeathSprout DeathSprout\nWindy Windy\nZhiyu\u0026#39;s Blog JUST DO IT ! ┏ (゜ω゜)= ☞ Zhonghao Sun Never thought of betrayal, nor need to speak of loyalty\nMiroier keep calm and carry on · GitHub Home\nZhangZhao A Lazy Programmer\nMakiras Although the sun shine, leave not your cloak at home.\nDamonZhang We are all in the gutter, but some of us are looking at the stars. · GitHub Home\nPYQ PYQ\u0026#39;s Blog | PYQ Blog\nQIN2DIM A creator from China.\ncodeslogan If not me, who?\nshiroha Narase Family\u0026#39;s General Store\n果壳儿 University of Chinese Academy of Sciences Community\nManim Kindergarten A group of manim Chinese users.\n苏剑林 Scientific Space - Aspiring to be a Little Peter Pan\nGuangzheng Li Reading and thinking, truth and freedom\nWeiyang Jin Current homepage of the SOTA paper rating author\nSam Altman AI is cool i guess\nWelcome to apply for a friend link~\nPlease leave a message in the following format\n1 2 3 4 站点标题: 小冰 站点简介: 小冰爱吃盐！ 博客地址: https://blog.bj-yan.top/ 头像地址: https://avatars.githubusercontent.com/u/44976445 ","permalink":"https://blog.bj-yan.top/en/friends/","summary":"\u003cblockquote\u003e\n\u003cp\u003eOccasionally I check in on my friends\u0026rsquo; friend links. If some old friends\u0026rsquo; domains or websites have become invalid, they may be temporarily removed. Please leave a message with your new site information~ (The following list is not ranked)\u003c/p\u003e\n\u003c/blockquote\u003e\n\u003chr\u003e\n\u003cdiv class=\"friend-link-div\"\u003e\n    \u003ca href=\"https://yuzhang.wang/\" title=\"YuZhangWang 的领域\" class=\"friend-link\" target=\"_blank\" rel=\"friend noopener noreferrer\"\u003e\n\n    \u003cdiv class=\"friend-link-avatar\"\u003e\n        \u003cimg src=\"https://gcore.jsdelivr.net/gh/YuZhangWang/Creative-pictures02@master/img/202210171416164.png\" class=\"friend-avatar\" loading=\"lazy\" decoding=\"async\" width=\"64\" height=\"64\" alt=\"YuZhangWang 的领域\" onerror=\"this.onerror=null;this.src='/friend_avator/default.svg'\"\u003e\n    \u003c/div\u003e\n    \u003cdiv class=\"friend-link-info\"\u003e\n        \u003ci class=\"fa fa-link\" aria-hidden=\"true\"\u003e\u003c/i\u003e\n        \u003ci class=\"friend-name\"\u003eYuZhangWang 的领域\u003c/i\u003e\n        \u003cp class=\"friend-bio\"\u003ePLUSULTRA!\u003c/p\u003e\n    \u003c/div\u003e\n    \u003c/a\u003e\n\u003c/div\u003e\n\n\u003cdiv class=\"friend-link-div\"\u003e\n    \u003ca href=\"https://russellwzr.github.io/\" title=\"Russell\" class=\"friend-link\" target=\"_blank\" rel=\"friend noopener noreferrer\"\u003e\n\n    \u003cdiv class=\"friend-link-avatar\"\u003e\n        \u003cimg src=\"/friend_avator/default.svg\" class=\"friend-avatar\" loading=\"lazy\" decoding=\"async\" width=\"64\" height=\"64\" alt=\"Russell\" onerror=\"this.onerror=null;this.src='/friend_avator/default.svg'\"\u003e\n    \u003c/div\u003e\n    \u003cdiv class=\"friend-link-info\"\u003e\n        \u003ci class=\"fa fa-link\" aria-hidden=\"true\"\u003e\u003c/i\u003e\n        \u003ci class=\"friend-name\"\u003eRussell\u003c/i\u003e\n        \u003cp class=\"friend-bio\"\u003eKeep eating codes!\u003c/p\u003e\n    \u003c/div\u003e\n    \u003c/a\u003e\n\u003c/div\u003e\n\n\u003cdiv class=\"friend-link-div\"\u003e\n    \u003ca href=\"https://bosswnx.xyz/\" title=\"Nelson Boss\" class=\"friend-link\" target=\"_blank\" rel=\"friend noopener noreferrer\"\u003e\n\n    \u003cdiv class=\"friend-link-avatar\"\u003e\n        \u003cimg src=\"https://avatars.githubusercontent.com/u/45135497?s=128\u0026amp;v=4\" class=\"friend-avatar\" loading=\"lazy\" decoding=\"async\" width=\"64\" height=\"64\" alt=\"Nelson Boss\" onerror=\"this.onerror=null;this.src='/friend_avator/default.svg'\"\u003e\n    \u003c/div\u003e\n    \u003cdiv class=\"friend-link-info\"\u003e\n        \u003ci class=\"fa fa-link\" aria-hidden=\"true\"\u003e\u003c/i\u003e\n        \u003ci class=\"friend-name\"\u003eNelson Boss\u003c/i\u003e\n        \u003cp class=\"friend-bio\"\u003eSwim until the sea turns blue\u003c/p\u003e\n    \u003c/div\u003e\n    \u003c/a\u003e\n\u003c/div\u003e\n\n\u003cdiv class=\"friend-link-div\"\u003e\n    \u003ca href=\"https://www.junyi.dev/\" title=\"Junyi\u0026#39;s Lab\" class=\"friend-link\" target=\"_blank\" rel=\"friend noopener noreferrer\"\u003e\n\n    \u003cdiv class=\"friend-link-avatar\"\u003e\n        \u003cimg src=\"https://avatars.githubusercontent.com/u/14367694\" class=\"friend-avatar\" loading=\"lazy\" decoding=\"async\" width=\"64\" height=\"64\" alt=\"Junyi\u0026#39;s Lab\" onerror=\"this.onerror=null;this.src='/friend_avator/default.svg'\"\u003e\n    \u003c/div\u003e\n    \u003cdiv class=\"friend-link-info\"\u003e\n        \u003ci class=\"fa fa-link\" aria-hidden=\"true\"\u003e\u003c/i\u003e\n        \u003ci class=\"friend-name\"\u003eJunyi\u0026#39;s Lab\u003c/i\u003e\n        \u003cp class=\"friend-bio\"\u003eA place for self-amusement\u003c/p\u003e","title":"Friends"}]