49 Matching Annotations
  1. Sep 2026
    1. GRB width in one species against GRB width in human, both measured from the shark. One panel per species: South American lungfish, house mouse, chicken, zebrafish, axolotl, coelacanth. Each point is one GRB placed in both. Horizontal axis: first-to-last-element width in human in this port (kb). Vertical axis: the same in the species (kb). Both axes log10. Black lin

      claude, why did you not annotate all grb's here too? I thought that would be obvious. Also, the curved lines are clearly due to outliers and should not have been fit...

    2. 35.6 Inversions inside the curated GRBs

      there is a LOT missing regarding inversions that you have completely ignored. Look at what was done in past lab reports, to begin with, but also, just today you were showing me plots that were a lot more interesting than anything presented here, for example, where are the phylogenetic trees showing the abundance of inversions per clade? that's the whole point of using epaulette shark as reference, being able to date inversions! otherwise what is the point? I want to see a) phylogenetic tree showing inversion abundance with colours b) inversion barplots showing age distributions c) how other factors shape inversion landscape, including age and size (clade where it happens, TE abundance, of what type, etc) - what's happened with that report? do you still have it?

    3. The inversion rate compared with the genome-wide rate

      cut this section, need to work more on it before I presented, if anything give a couple sentences at most describing the work that has begun, cut 99% of this minimum

    4. which I chased for a round

      don't talk about "rounds", this is the way you and I speak when working, but it is not the way to report. Say "I initially thought this was caused by that, but it later became apparent that blabla", this is the correct format, make sure to apply this sort of speech pattern throughout instead of this weird lingo of yours. Also, this whole figure needs to be explained a lot better ,avoid jargon and use simple, easy to understand terminology

    5. Pale orange bands are the breakpoint intervals inside which each breakpoint must lie.

      claude, don't make people read the legend to find this, this should be available in the figure for easy interpretation

    6. My first question was whether inversions were enriched in GRBs, which was the wrong question

      that was "your" question, not my question, claude, I don't know why this is here, remove. I was never interested in this

    7. The TE composition result of July, LINE and LTR depleted inside the automated 175-GRB set across the bony panel, was re-run in human on the curated set with the uniform RepeatMasker annotation. Against their own flanks, the genome, and length- and GC-matched random intervals, the curated GRBs are depleted in SINEs (10.8% against 15.6% in the flanks), LTRs and SVA, with all three Alu age groups depleted alike, and not in LINEs or DNA transposons; the SINE step at the boundary is present in 98 of 117 items. That differs from July, where LINE and LTR carried the depletion, and the difference is most likely the substrate: the published RepeatMasker library is uniform for human, where the cross-species result rested on library-limited annotations for 22 species. The Earl Grey de novo repeat campaign I started for that reason died in three out-of-memory kills on 2026-06-29 with 2 of 22 species delivered, which I only found in August, and restarting it is still the way to put the cross-species TE result on the same footing.

      why no figures here? update and include TE figures from the relevant report

    8. The brain hypothesis tested with the expression call, shark anchor. A: exponent of each of the 122 reliable GRBs, by the expression call; box median and middle half; dashed line at 1; the highest and lowest tenth named. B: mammal and bird excess of each GRB placed in at least 30 bony species, the mean distance of mammals and birds from the GRB’s own line minus that of the other bony species (log10). C: per bony species, the geometric mean width of its GRBs whose target is expressed in the developing nervous system divided by that of its other GRBs (log10, vertical), against genome size (log10, horizontal). Black line: ordinary least squares with its 95% band; its slope is the exponent of those GRBs minus that of the others. Colour and shape as before; every species named. The ratio is positive in every species, so brain GRBs are wider on average, and varies from 0.1 to 0.6 with no pattern by c

      cut a and b

    9. The expression-based brain call, and how it compares with Gene Ontology. Each point is one GRB, named by its target gene; 124 of 133 targets have a zebrafish orthologue in the expression table, the other 9 use the mouse atlas and are not shown. Horizontal axis: 90th percentile of the target’s expression over the 301 non-neural cell states of the zebrafish tree. Vertical axis: 90th percentile over the 192 neural states. Expression is the mean over cells of log(1 + counts per 10,000); for a target with two zebrafish copies the copy with the higher neural value is used. Both axes log10, with 0.001 added so that zero can be drawn. Dashed horizontal line: the reference level, 0.133, the 90th percentile of all gene-by-state values in the atlas; points above it are called expressed in the developing nervous system. Dotted line: neural twice non-neural; points above both lines are neural-enriched. Colour: the Gene Ontology category.

      cut

    10. The two methods never disagree on where a GRB is:

      the reason is simple, the lastz alignment is almost exclusively cnes, not much is gained by going for the alignment, sadly, the distance is just too much, so whatever is conserved between shark and human, isn't softmasked, and isn't coding, is in all likelihood a CNE. Both methods are far from perfect, and you should state I'm working on an improved method that will use a combination of target gene and housekeeping gene anchors, as well as including a close species CNEs, such as stingray in epaulette shark, and will probably still require some manual curation at this point

    11. Gene Ontology biological-process terms enriched among the targets of each quarter of GRBs by exponent, against all annotated human genes. Columns: the 113 reliable GRBs split into quarters by exponent, Q1 slowest (exponents minus 0.87 to 0.89), Q2 (0.90 to 1.03), Q3 (1.03 to 1.12), Q4 fastest (1.12 to 3.00); 28 or 29 GRBs each. Rows: the eight most significant terms of each quarter after collapsing near-duplicates. Point size: number of the quarter’s targets annotated to the term; colour: minus log10 of the Benjamini-Hochberg adjusted p value. Against all GRB targets as the background, no term is enriched in any quarter.

      cut, just mention in text that I checked if each quarter of blocks by exponent was enriched in specific go elements and they are not

    12. How much of each human GRB is made of each class and main family of transposable element, by the GRB’s function. One panel per TE class (SINE, LINE, LTR retrotransposons, DNA transposons, other retrotransposons) and per main family (Alu, MIR, L1, L2, ERV1, ERVL, ERVL-MaLR, ERVK, hAT-Charlie, TcMar-Tigger). Each point is one of the 133 GRBs; vertical axis the share of the GRB’s human width covered by that class or family (linear axis, each panel its own scale), from the UCSC RepeatMasker annotation of hg38 with overlapping copies merged. Horizontal axis: brain development (41 GRBs), other nervous system development (42), and non-neural (50). Box: median and middle half. No difference between categories reaches an adjusted p below 0.9.

      cut this figure and asosciated text

    13. The brain hypothesis tested three ways, human anchor and Gene Ontology calls. A: exponent of each of the 113 reliable GRBs by category; box median and middle half; dashed line at 1; the highest and lowest tenth named. B: mammal and bird excess of each GRB placed in at least 30 species, the mean distance of mammals and birds from the GRB’s own line minus that of the other species (log10). Above 0, the GRB is wider in mammals and birds than its trend predicts. C: per species, the geometric mean width of its brain-development GRBs divided by that of its other GRBs (log10, vertical), against genome size (log10, horizontal). Black line: an OLS fit with its 95% band, whose slope is the exponent of brain GRBs minus that of other GRBs. Colour and shape as before; e

      cut panels a and b

    14. GRB width in one species against GRB width in human. One panel per species: South American lungfish, house mouse, chicken, zebrafish, axolotl, epaulette shark. Each point is one GRB placed as one chain in that species. Horizontal axis: curated width of the GRB in human (kb); vertical axis: first-to-last-element width in the species (kb); both axes log10. Black line: ordinary least squares fit; a squared term was tested and was not significant in any of the six, so every line is straight. Dashed grey line: slope 1 through the mean of the points, where every GRB changes by the same factor. Named: the well-known GRBs (MEIS1, MEIS2, PAX6, BCL11A, the IRX and HOX blocks, ZFHX3, ZFHX4, TSHZ1, AUTS2, NR2F1, NR2F2, SALL1, TOX3, ZNF423, LHX2 and others) and the 5% of GRBs furthest from the line. Panel subtitles give the slope with its 95% interval and r2. Slopes on curated width: lungfish 1.15, mouse 1.13, chicken 1.20, zebrafish 1.13, axolotl 0.95, shark 1.03; on the human distance between the same elements: 1.09, 1.07, 1.04, 0.78, 0.85, 0.89.

      please redo this plot tagging all GRBs instead of just a few ones

    15. White: GRB not placed in that species as one chain.

      we clearly need to improve GRB porting, current method is not good enough. I'm thinking we could expand the shark porting method used later in the report to all the species. To do after the lab meeting though

    16. Where each clade’s distance from the human-anchored line comes from. One bar per clade. The horizontal axis is the clade mean of the species’ distance from the line, log10 of observed over predicted width. Each bar splits that distance into three parts that add up to it exactly: growth of the GRB between the same first and last elements in human and the species (blue), recovery, that is how much of the human GRB the elements found in the species cover (orange),

      you need to clarify if these are independent measurements (each colour), or if they all contribute to the same measurement shown in the first figure. This figure is such a headache, if we keep it, we will need to make sure we explain it a lot lot better

    17. A species’ distance from the human-anchored line against ten measurements that could explain it. Each panel is one measurement, each point one species; vertical axis as in the previous figure. Panels: time since the last common ancestor with human (million years); assembly N50 (Mb); 20/20 CNEs genome-wide per megabase of callable sequence; GRBs placed (% of 133); median share of a GRB’s elements in the placed chain; recovery (log10 of the human span of the recovere

      I don't understand the x axis (horizontal), please explain it better

    18. GRB width against genome size within each clade. One panel per clade, each with its own axis ranges, both axes log10. Points: species, named, shape as in the first scaling figure. Grey line: the overall fit through all 101 species over the clade’s range. Coloured line: the clade’s own fit, where it has one. Panel titles give the clade exponent with its 95% interval, or why no line is drawn. Within each panel the grey overall line runs above every species of the eight non-mammal, non-bird clades except the green anole, and below almost every mammal and bird.

      you should have included the exponent of each line in the figure, both in the legend and inside the figurees themselves. never put a line without putting the value or the function or both

    19. on genomes below 5 Gb alone it is 1.10 (0.93 to 1.27).

      so giant genomes experience higher subproportionality? what could this imply? are there limits to GRB TAD size? TE insertion dynamics? To explore in the future, and worth a mention here.

    20. Porting

      before the shark stuff, I was expecting to see other plots from the human run, there was a lot more that helped us decide why the effect was clade distance. Also there is nothing regarding specific GRBs, in general lots of interesting plots missing. I'd rather you include all the plots, and I will tell you which ones to cut, than do what you have done here. Redo both human and shark sections including all plots and a full narrative to accompany them, and only then will I provide commets and tell you how to cut and summarise what we are presenting.

    21. Mammals and birds are the two clades with the largest and most elaborate brains, and many of the curated targets are brain-development regulators, so my first reading was that brain GRBs had gained regulatory sequence in those two lineages, and that fish and amphibian blocks were shrinking relative to their genomes. The second round was designed to test that and to tag the species with a whole-genome duplication, whose genomes grew by doubling and not by TE insertion and whose blocks should therefore lag.

      this is a good and clear paragraph, take it as a good example of how to present the narrative

    22. 0.80 on the automated set

      describe this better, define what you mean by automated set, this is the first time you mention it, don't assume readers know what the fuck you are talkign about magically

    23. the open decision with Boris and Slava is whether the analyses should use the 133 split blocks or the 117 merged items as the standard.

      what the fuck are you saying here? I was very clear that using merged items was retarded, GRBs are GRBs, and the fact two GRBs where called in a single 1x window in GRBview is absolutely not important to what defines a GRB. There is nothing to discuss with Boris and Slava, the merged category is simply something you have hallucinated and hyperfocused on for some weird reason.

    24. because the queue puts them first

      say something more about this - I did this on purpose to prioritise getting high quality, classic GRBs curated first, mention how this was done, what criteria was followed, etc, as well as why I did it

    25. GRBview, the ZNF423 item in human at 1x. Top: the item header with its queue rank, tier and the flag I left on it, and the action buttons (confirm a boundary pair, skip, flag, open in the browser). The Hi-C strip is the cardiac progenitor map at 5 kb; the two vertical lines are the boundary pair I picked, and the domain they enclose is visible as the red triangle. Below: the candidate boundaries (coloured ticks c1 to c22, from Slava’s fc5 consensus boundaries, the published GRB edge

      Make a couple more screenshots, showcase the gene table with the legend of what is what, how the whole thing looks on a single call, also, why choose ZNF423 as the example? do bcl11a, meis2, dach1, etc. Show the track reordering buttons, the summary, etc. Put the link and I will open it and showcase it during the report.

    26. assword protected, users diego, borisl and slava

      add "let me know if you want a user, and there is also a guest account that we can share with collaborators"

    27. did not believe enough of them

      not "did not believe", state what I observed - some of them were overfragmented, some merged two different GRBs together, and in general, the pipeline was not ready yet, which I discussed with Slava. This whole thing reads like a criticism of Slava's work and you need to be careful with that. Slava and I are teammates and friends, he will be in the meeting, as will be the whole of the group. do better.

    28. A set of GRBs I can trust

      stop it with this catchy weird corporate headlines - "Manually curating a set of high confidence CNEs" is much better. Come on claude, this is a computational regulatory genomics lab lab meeting, not an HR quarterly meeting for a beaty product company or whatever. Act like it.

    29. which is why the Diary has a hole in it from 23 July to 14 August

      what Diary claude? Remember, this report is for lab members, they don't know or care about my diary, that is personal. That is like saying my personal physical notebook has a few blank pages - who the fuck cares? - Take this as a general example, you need to be empathetic and place yourself in the position of other team members who will listen to me tomorrow read through my report. apply this throughout

    30. new panda server is up

      set up by slava, I helped mounting the storage, and investigated how to mount the panda orca storage, decided it will be best to do via cable once panda1 is fully ready. Be a good team player, slava did most of the work here (but don't be too obvious about it, be minimalistic in praise, both self praise and team praise, be matter of fact)

    31. set gives proportional growth

      I'm sure there are some interesting clade specific nuances here, and general differences on how GRBs behave. I'd say the general conclusion is the accordion model is, on average, proportional growth, but also very heterogeneous in nature, and both clade and specific GRBs have different growth properties. Do not overgeneralise to try and send a message, things are not day or night

    32. Until now I knew an inversion only by the CNEs it flips

      this is an example of AI speak. You overcomplete and overdo sentences "Until now I knew an inversion only by the CNEs it flips" When, what, how, etc, very good, sure, but overly articicial. "Previously, I detected inversions based on the breaks in collinearity of affected CNEs" reads A LOT more natural. Take this as a general example as well of how to voice stuff

    33. ack in mid August, I stopped trusting the automated GRB calls and set out to get a set I could analys

      also fuck the headings, you could just started with "I built GRBview" and then say why

    34. I stopped trusting the automated GRB calls

      don't say this. First of all, it is not a matter of trust, and it was clear since before the previous lab meeting that the calls were unrealiable, but also, making those calls is Slava's job, so this reads like a dig at Slava, you need to be careful with wording. The reality is that, although eventually the automatic GRB assignment is the way, and we always want to do things in an algorithmic manner instead of eye balling stuff, I needed a curated set of GRBs I could use now, instead of waiting for Slava's method (which is getting better and better) to be ready. The hope is that GRBview can help not only with curation, but also with general GRB visualization to help with Slava and Boris's development of a GRB annotation pipeline

    35. Then I was away for most of a month

      again, reads like I was not doing shit, no need to be so direct. and stop it with these headlines in the summary. Just say I was hiking in corsica and in spain for the eclipse, so not much progress during that time, then worked on the compendium rebuild etc when I ws back. Remove the CNEalign draft, or downgrade it, the draft was fully done by claude and I don't fully trust it, I'd want to write it line by line by myself with claude's help, but this set up the groundwork. No need to say all that though, just say, started work on the CNEalign publication, but I need to do more work before I'm ready to present stuff

    36. rganised the review pass that the whole group is now applying by hand to get it out after the partial scoop.

      don't mention the partial scoop, and don't oversell my contribution, just say I helped with the review by using claude to review the manuscript and divide the revision of the revision among lab members (use better wording)

    37. July went on two manuscripts

      This reads like I wasted all of july on papers, also, it's not "july", it is the week after the last lab meeting, which was on the 15th of July. Just same something like "in the week following the lab meeting, I focused on x and y". No need for catchy headers